Cross-Species Aging Knowledge Integration into Agentic AI Platform Uncovers Conserved Mechanisms
The 21 matches
- [1] § Material and Methods › Data Source Curation and Collection ↔ pipeline/11_ploting/all_figures-raidhani.ipynb, lines 5255–5312 · score 1.00 · Worm Interactome Database, BioGrakn, CROssBAR, DTINet, BindingDB, BioGRID
- [2] § Material and Methods › Data Source Curation and Collection ↔ Frontend/components/AboutUs.py, lines 187–246 · score 0.96 · CellAge, GeneAge, MetaboAge, AgeAnno, DrugAge, AgeXtend
- [3] § Material and Methods › Knowledge Graph Embedding (KGE) Training ↔ Backend/dgl-ke/python/dglke/utils.py, lines 252–350 · score 0.85 · SimplE, ComplEx, TransE, RotatE, loss function, DGL KE
- [4] § Results › EvoAge Enables Accurate Hypothesis Generation and Experimental Validation ↔ Frontend/hypo_agents.py, lines 395–452 · score 0.81 · Amyloid Precursor Protein, secretase enzyme, amyloidogenic processing, APP, compartmental, Postsynapse
- [5] § Material and Methods › Knowledge Graph Embedding (KGE) Training ↔ Backend/dgl-ke/python/dglke/eval.py, lines 41–105 · score 0.80 · SimplE, ComplEx, TransE, RotatE, DGL KE, RESCAL
- [6] § Results › A Systems-Scale Knowledge Graph Integrating Aging Biology Across Model Species ↔ pipeline/02_data_processing/drugage.ipynb, lines 1–65 · score 0.79 · Saccharomyces cerevisiae, Mus musculus, Danio rerio, Caenorhabditis elegans, Drosophila melanogaster, pipeline
- [7] § Results › A Systems-Scale Knowledge Graph Integrating Aging Biology Across Model Species ↔ pipeline/04_orthology_mapping/Run_2_Final_map_orthologs_to_human_desc.R, lines 1–41 · score 0.79 · Saccharomyces cerevisiae, Mus musculus, Danio rerio, Caenorhabditis elegans, Drosophila melanogaster, pipeline
- [8] § Material and Methods › Data Preprocessing and Harmonization ↔ pipeline/02_data_processing/tarkg.ipynb, lines 1–97 · score 0.78 · NCBI Gene, UniProt, PubChem, Cellular Component, Reactome, HPO
- [9] § Material and Methods › Backend API and Service Architecture ↔ Frontend/evo-utils/src/kani_utils/kani_streamlit_server.py, lines 147–276 · score 0.78 · get_sample_triples, check_relationship, search_biological_entities, server, subgraph, retrieval
- [10] § Material and Methods › Backend API and Service Architecture ↔ Frontend/components/MicroServices.py, lines 651–716 · score 0.77 · entity_relationships, check_relationship, sample triples, service, subgraph, Link Prediction
- [11] § Material and Methods › Data Preprocessing and Harmonization ↔ pipeline/02_data_processing/PrimeKG.ipynb, lines 1–90 · score 0.77 · NCBI Gene, DrugBank, PubChem, Cellular Component, MONDO, HPO
- [12] § Results › Systematic Optimization of Knowledge Graph Embeddings Reveals Architectural Dependencies ↔ Backend/dgl-ke/python/dglke/utils.py, lines 252–350 · score 0.72 · SimplE, ComplEx, TransE, RotatE, loss, RESCAL
- [13] § Results › Systematic Optimization of Knowledge Graph Embeddings Reveals Architectural Dependencies ↔ pipeline/11_ploting/all_figures-raidhani.ipynb, lines 1679–1745 · score 0.71 · SimplE, ComplEx, TransE, RotatE, MRR, hit
- [14] § Results › A Systems-Scale Knowledge Graph Integrating Aging Biology Across Model Species ↔ Frontend/components/LoginSignup.py, lines 208–282 · score 0.68 · human centric, model organisms, cross species, unified, harmonization, EvoAge
- [15] § Results › A Conversational AI Interface Leverages the EvoAge KG for Discovery and Validation ↔ Frontend/components/AboutUs.py, lines 187–246 · score 0.68 · FastAPI, Neo4j, capabilities, powered, Streamlit, KE
- [16] § Material and Methods › Orthology-Based Cross-Species Integration ↔ pipeline/04_orthology_mapping/Run_2_Final_map_orthologs_to_human_desc.R, lines 1–41 · score 0.58 · human ortholog mapping, Ensembl, cerevisiae, rerio, musculus, melanogaster
- [17] § Results › A Systems-Scale Knowledge Graph Integrating Aging Biology Across Model Species ↔ Frontend/components/LoginSignup.py, lines 208–282 · score 0.58 · biological networks, human centric, connectivity, discovery, EvoAge, Biology
- [18] § Results › A Conversational AI Interface Leverages the EvoAge KG for Discovery and Validation ↔ Frontend/agents.py, lines 144–204 · score 0.58 · Search Biological Entity, triple score, invokes, Agent, validates, ranked
- [19] § Material and Methods › Graph Database Implementation ↔ pipeline/09_evoage_vs_other/evoage_vs_biochat_escargot/Biochatter/generate_schema_info.py, lines 46–137 · score 0.55 · graph database, Neo4j, APOC, Cypher, Knowledge Graph, edges
- [20] § Results › A Systems-Scale Knowledge Graph Integrating Aging Biology Across Model Species ↔ pipeline/02_data_processing/ageannomo.ipynb, lines 171–232 · score 0.54 · Danio rerio, Drosophila melanogaster, genomic, schema, mouse, nodes
- [21] § Results › EvoAge Enables Accurate Hypothesis Generation and Experimental Validation ↔ Frontend/hypo_agents.py, lines 395–452 · score 0.50 · amyloidogenic processing, endocytic, secretase, postsynaptic, BACE1, phenotype
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 5,312 lines · 199 KB · MIT · 2 matches
- # %%
- import pandas as pd
- # %% [markdown]
- # ## Main Figure 1
- # %% [markdown]
- # ## 1b
- # %%
- """
- EvoAge KG — Figure 1b: single-ring sunburst/donut of data sources,
- wedge size = log10(Total edges), colored by KG_Type category.
- Reads figure1b.csv (columns: Source, KG_Type, Total, log) and reproduces
- evoage_sunburst_1b.svg. Output SVG text stays fully editable (svg.fonttype='none').
- """
- import pandas as pd
- import matplotlib.pyplot as plt
- import matplotlib as mpl
- mpl.rcParams['font.family'] = 'DejaVu Sans'
- # ----------------------------------------------------------------------
- # Editable-text SVG output (Rai's standing preference)
- # ----------------------------------------------------------------------
- mpl.rcParams['svg.fonttype'] = 'none'
- # ----------------------------------------------------------------------
- # Config
- # ----------------------------------------------------------------------
- CSV_PATH = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/figure1b.csv" # update path as needed, e.g. server path
- OUT_PATH = "/storage/Arushi/090526_EvoAge/kg_formation/all_figures/FIG1/evoage_sunburst_1b_regen.svg"
- CATEGORY_ORDER = ["Aging", "Biomedical", "Species connection"]
- CATEGORY_COLORS = {
- "Aging": "#e8b4b8", # salmon/pink
- "Biomedical": "#7fb3b8", # teal
- "Species connection": "#f1c88b", # gold
- }
- DONUT_WIDTH = 0.55 # ring thickness as fraction of radius (donut hole size)
- WEDGE_EDGE_COLOR = "white"
- WEDGE_EDGE_WIDTH = 1.0
- START_ANGLE = 90 # start at 12 o'clock
- CLOCKWISE = True # counterclockwise=False in pie()
- SHOW_RAW_NUMBERS = True # print each source's raw Total on its wedge
- RAW_LABEL_FONTSIZE = 6
- RAW_LABEL_MIN_PCT = 0.0 # raise this (e.g. 0.5) to hide labels on tiny wedges
- def format_raw(x: float) -> str:
- """Human-readable raw count, e.g. 1055910438 -> '1.06B'."""
- if x >= 1e9:
- return f"{x/1e9:.2f}B"
- if x >= 1e6:
- return f"{x/1e6:.1f}M"
- if x >= 1e3:
- return f"{x/1e3:.1f}K"
- return f"{int(x)}"
- # ----------------------------------------------------------------------
- # Load + prep data
- # ----------------------------------------------------------------------
- df = pd.read_csv(CSV_PATH)
- # clean thousands-separator formatted totals (e.g. "4,71,00,000")
- df["Total"] = df["Total"].astype(str).str.replace(",", "").astype(float)
- # category order = alphabetical (Aging, Biomedical, Species connection);
- # within each category, rows are already sorted descending by 'log' —
- # preserve that order rather than re-sorting, so ties/order match the source file
- df["KG_Type"] = pd.Categorical(df["KG_Type"], categories=CATEGORY_ORDER, ordered=True)
- df = df.sort_values(["KG_Type"], kind="stable").reset_index(drop=True)
- sizes = df["log"].values # <-- wedge size uses log values
- colors = df["KG_Type"].map(CATEGORY_COLORS).values
- grand_total = df["Total"].sum()
- # ----------------------------------------------------------------------
- # Plot
- # ----------------------------------------------------------------------
- fig, ax = plt.subplots(figsize=(9.57, 11.27)) # matches ~689x811pt canvas
- # autopct only receives a percentage, so use a row counter (pie() calls it
- # once per wedge, in the same order as `sizes`/`df`) to map back to the
- # actual raw Total for that source.
- raw_values = df["Total"].tolist()
- _row = {"i": 0}
- def raw_number_label(pct):
- i = _row["i"]
- _row["i"] += 1
- if not SHOW_RAW_NUMBERS or pct < RAW_LABEL_MIN_PCT:
- return ""
- return format_raw(raw_values[i])
- wedges, _, raw_labels = ax.pie(
- sizes,
- colors=colors,
- startangle=START_ANGLE,
- counterclock=not CLOCKWISE,
- radius=1.0,
- wedgeprops=dict(
- width=DONUT_WIDTH,
- edgecolor=WEDGE_EDGE_COLOR,
- linewidth=WEDGE_EDGE_WIDTH,
- ),
- autopct=raw_number_label,
- pctdistance=1.0 - DONUT_WIDTH / 2, # center of the donut ring
- )
- for lbl in raw_labels:
- lbl.set_fontsize(RAW_LABEL_FONTSIZE)
- lbl.set_ha("center")
- lbl.set_va("center")
- ax.set_aspect("equal")
- # ----------------------------------------------------------------------
- # Center annotation
- # ----------------------------------------------------------------------
- ax.text(
- 0, 0.12, "1",
- ha="center", va="center", fontsize=22, fontweight="bold",
- )
- ax.text(
- 0, -0.02, f"~{grand_total/1e9:.2f} billion",
- ha="center", va="center", fontsize=13,
- )
- ax.text(
- 0, -0.14, "head ←|||◀ tail",
- ha="center", va="center", fontsize=11,
- )
- # ----------------------------------------------------------------------
- # Title + legend
- # ----------------------------------------------------------------------
- ax.set_title("EvoAge", fontsize=18, pad=20)
- legend_handles = [
- plt.Rectangle((0, 0), 1, 1, color=CATEGORY_COLORS[cat])
- for cat in CATEGORY_ORDER
- ]
- ax.legend(
- legend_handles,
- CATEGORY_ORDER,
- loc="upper center",
- bbox_to_anchor=(0.5, 0.02),
- ncol=3,
- frameon=False,
- fontsize=10,
- )
- fig.tight_layout()
- fig.savefig(OUT_PATH, format="svg")
- print(f"Saved: {OUT_PATH}")
- # %% [markdown]
- # ## 1c
- # %%
- """
- Vertical grouped bar charts comparing gene counts at the "121" vs
- "121_12M" ortholog-mapping stage, per species, for each KG resource.
- Plot 1: Aging KG vs Biomedical KG
- Plot 2: Aging KG vs EvoAge KG vs Biomedical KG
- Input : combined_gene_counts.csv
- columns -> Species, Aging_121, Aging_121_12M,
- Biomedical_121, Biomedical_121_12M,
- EvoAge_121, EvoAge_121_12M
- Output: SVG (+PNG) files, vector, publication-ready.
- """
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- from matplotlib.patches import Patch
- # ----------------------------------------------------------------------
- # 0. Paths -- EDIT THESE to match your local setup
- # ----------------------------------------------------------------------
- INPUT_CSV = "combined_gene_counts_no_human.csv"
- OUT_DIR = "." # where the SVG/PNG figures will be written
- # ----------------------------------------------------------------------
- # 1. Style
- # ----------------------------------------------------------------------
- plt.rcParams.update({
- "font.family": "sans-serif",
- "font.size": 11,
- "axes.spines.top": False,
- "axes.spines.right": False,
- "svg.fonttype": "none", # keep text editable in Illustrator/Inkscape
- })
- COLOR_121 = "#3F6FA6" # blue -> "121" (before)
- COLOR_121_12M = "#C1440E" # orange -> "121_12M" (after)
- # established dataset accent colors (used for the small dataset tick labels)
- DATASET_COLORS = {
- "Aging": "#F08080",
- "Biomedical": "#E8A33D",
- "EvoAge": "#4C9AAE",
- }
- def plot_gene_count_bars(df, datasets, title, outfile,
- bar_width=0.38, ds_gap=0.25, group_gap=1.1,
- figsize_per_col=1.15):
- """
- df : the combined_gene_counts dataframe
- datasets : list of prefixes, e.g. ["Aging", "Biomedical"]
- or ["Aging", "EvoAge", "Biomedical"]
- Each dataset contributes one pair of bars (121 / 121_12M) per species.
- """
- species_list = df["Species"].tolist()
- n_ds = len(datasets)
- fig_w = max(7, figsize_per_col * n_ds * len(species_list) + 2)
- fig, ax = plt.subplots(figsize=(fig_w, 6.2))
- x_cursor = 0.0
- group_centers, group_labels = [], []
- tick_pos, tick_labels, tick_colors = [], [], []
- for sp in species_list:
- row = df[df["Species"] == sp].iloc[0]
- group_start = x_cursor
- for ds in datasets:
- y0 = row[f"{ds}_121"]
- y1 = row[f"{ds}_121_12M"]
- x_center = x_cursor
- if pd.notna(y0) and pd.notna(y1):
- ax.bar(x_center - bar_width / 2, y0, width=bar_width,
- color=COLOR_121, edgecolor="white", linewidth=0.6, zorder=3)
- ax.bar(x_center + bar_width / 2, y1, width=bar_width,
- color=COLOR_121_12M, edgecolor="white", linewidth=0.6, zorder=3)
- else:
- ax.text(x_center, ax.get_ylim()[1] if ax.get_ylim()[1] > 0 else 1000,
- "n/a", ha="center", va="bottom", fontsize=8,
- color="grey", style="italic")
- tick_pos.append(x_center)
- tick_labels.append(ds)
- tick_colors.append(DATASET_COLORS.get(ds, "black"))
- x_cursor += (bar_width * 2 + ds_gap)
- group_centers.append((group_start + x_cursor - (bar_width * 2 + ds_gap)) / 2)
- group_labels.append(sp)
- # light vertical separator between species groups
- if sp != species_list[-1]:
- sep_x = x_cursor - ds_gap / 2 + group_gap / 2
- ax.axvline(sep_x, color="#DDDDDD", lw=0.8, zorder=0)
- x_cursor += group_gap
- # dataset-level ticks (small, colored)
- ax.set_xticks(tick_pos)
- ax.set_xticklabels(tick_labels, rotation=90 if n_ds > 2 else 0, fontsize=8)
- for tick, c in zip(ax.get_xticklabels(), tick_colors):
- tick.set_color(c)
- # species-level labels below the dataset ticks
- for xc, sp in zip(group_centers, group_labels):
- ax.text(xc, -0.13, sp, transform=ax.get_xaxis_transform(),
- ha="center", va="top", fontsize=11.5, fontweight="bold")
- ax.set_xlim(-bar_width * 1.5, x_cursor - group_gap + bar_width * 1.5)
- ax.set_ylabel("Gene count", fontsize=12)
- ax.set_title(title, fontsize=13.5, fontweight="bold", pad=14)
- ax.margins(y=0.08)
- ax.grid(axis="y", color="#EAEAEA", lw=0.8, zorder=0)
- legend_elems = [
- Patch(facecolor=COLOR_121, label="121"),
- Patch(facecolor=COLOR_121_12M, label="121_12M"),
- ]
- ax.legend(handles=legend_elems, frameon=False, loc="upper left",
- bbox_to_anchor=(1.01, 1.0), title="Ortholog stage")
- fig.tight_layout()
- fig.savefig(f"{OUT_DIR}/{outfile}.svg", bbox_inches="tight")
- fig.savefig(f"{OUT_DIR}/{outfile}.png", dpi=300, bbox_inches="tight")
- plt.show()
- plt.close(fig)
- print(f"Saved {outfile}.svg / .png")
- if __name__ == "__main__":
- df = pd.read_csv(INPUT_CSV)
- # Plot 1: Aging vs Biomedical
- plot_gene_count_bars(
- df,
- datasets=["Aging", "Biomedical"],
- title="Gene Count Comparison: 1-to-1 vs 1-to-Many orthologs (Aging vs Biomedical KG)",
- outfile="genecount_bar_aging_biomedical",
- )
- # Plot 2: Aging vs EvoAge vs Biomedical
- plot_gene_count_bars(
- df,
- datasets=["Aging", "EvoAge", "Biomedical"],
- title="Gene Count Comparison: 1-to1 vs 1-to-1+1-to-M (Aging vs EvoAge vs Biomedical KG)",
- outfile="genecount_bar_aging_evoage_biomedical",
- )
- # %% [markdown]
- # ## 1d
- # %%
- import plotly.io as pio
- pio.renderers.default = "iframe"
- # %%
- """
- KG Node-Type Hierarchy Visualization (Sunburst ONLY)
- =====================================================
- Builds a multi-layer sunburst chart (Dataset -> Species -> Node type)
- for your Aging vs Biomedical knowledge graphs, for the node scheme:
- - Node (1:1 + 1:N)
- Chart details:
- - Two top-level branches: Aging (coral/red family) vs Biomedical
- (teal family -- same teal used previously for EvoAge) -- children
- inherit shades of their branch color.
- - Wedge SIZE is driven by log10(count+1), so small node types (e.g.
- Chemical, Protein) stay readable next to huge ones (PMID, Mutation)
- instead of being squeezed to an invisible sliver.
- - Labels and hover text still show the REAL, un-logged count (e.g.
- "Chemical 434,768"), not the log value -- log scale only affects size.
- - Combines all species into one figure per scheme (species is just one
- ring/layer of the hierarchy, so you still see per-species breakdown
- without needing six separate panels).
- Requirements:
- pip install pandas plotly kaleido
- SVG export (kaleido) note: recent kaleido versions (v1+) need a local
- Chrome install. If `write_image` fails, run this once in your terminal:
- plotly_get_chrome
- or pin an older, self-contained kaleido that needs no Chrome:
- pip install "kaleido==0.2.1"
- HOW TO USE
- ----------
- 1. Edit the FILES dict below so the paths point to your own CSVs.
- 2. Run: python kg_hierarchy_sunburst_aging_biomedical.py
- 3. Each figure is saved directly as a .svg file next to this script.
- (No HTML files are written -- SVG only, per your request.)
- """
- import numpy as np
- import pandas as pd
- import plotly.express as px
- # ---------------------------------------------------------------------------
- # 1. EDIT THESE PATHS to point at your own CSV files
- # ---------------------------------------------------------------------------
- FILES = {
- "1to1_1toN": {
- "Aging": "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/aging_121_12M_Species_NodeType_unique_counts.csv",
- "Biomedical": "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/Species_Biomedical_121_12M_NodeType_unique_counts.csv",
- },
- }
- SCHEME_TITLE = {
- "1to1_1toN": "Node (1:1 + 1:N)",
- }
- # Node-type row names (as they appear in the CSV) -> short display labels
- NODE_MAP = {
- "CellularComponent": "Cellular",
- "Disease": "Disease",
- "Phenotype": "Phenotype",
- "BiologicalProcess": "Biological",
- "PlantSpecies": "Plant",
- "MolecularFunction": "Molecular",
- "Pathway": "Pathway",
- "Mutation": "Mutation",
- "ChemicalEntity": "Chemical",
- "Protein": "Protein",
- "Mirna": "miRNA",
- "AnatomicalEntity": "Anatomy",
- "Gene": "Gene",
- "Tissue": "Tissue",
- "PMID": "PMID",
- }
- # Distinct color families for the two datasets (kept for reference/legend)
- # NOTE: Biomedical keeps the SAME teal family previously used for EvoAge.
- DATASET_COLORS = {
- "Aging": "#E85C6B", # coral/red family
- "Biomedical": "#1F7A8C", # teal family (same as former EvoAge)
- }
- # Distinct designated color per node type (leaf-level segments)
- NODE_COLORS = {
- "Cellular": "#F5E1A4",
- "Disease": "#B79FCB",
- "Phenotype": "#E0559B",
- "Biological": "#7FB8D6",
- "Plant": "#5B2A6E",
- "Molecular": "#D98E3F",
- "Pathway": "#7A4B32",
- "Mutation": "#9B9B9B",
- "Chemical": "#C9B458",
- "Protein": "#F0BFC8",
- "miRNA": "#8E4E8F",
- "Anatomy": "#E08D6D",
- "Gene": "#4472A8",
- "Tissue": "#E8A96B",
- "PMID": "#B7C88A",
- }
- # ---------------------------------------------------------------------------
- # 2. Reshape each CSV (wide: NodeType x Species) into long hierarchical rows
- # ---------------------------------------------------------------------------
- def load_long(dataset_name, path):
- df = pd.read_csv(path)
- df = df[df["NodeType"].isin(NODE_MAP.keys())].copy()
- df["NodeType"] = df["NodeType"].map(NODE_MAP)
- species_cols = [c for c in df.columns if c != "NodeType"]
- long_df = df.melt(id_vars="NodeType", value_vars=species_cols,
- var_name="Species", value_name="Count")
- long_df["Dataset"] = dataset_name
- long_df = long_df[long_df["Count"] > 0] # drop empty branches (cleaner chart)
- long_df["LogCount"] = np.log10(long_df["Count"] + 1) # drives wedge SIZE only
- return long_df[["Dataset", "Species", "NodeType", "Count", "LogCount"]]
- def build_hierarchy(scheme_key):
- files = FILES[scheme_key]
- frames = [load_long(ds_name, path) for ds_name, path in files.items()]
- return pd.concat(frames, ignore_index=True)
- def apply_node_colors(fig):
- """Give each NodeType leaf its designated color; Dataset/Species parent
- rings get a neutral gray so they don't inherit a random default color."""
- colors = [NODE_COLORS.get(lab, "#EDEDED") for lab in fig.data[0].labels]
- fig.data[0].marker.colors = colors
- return fig
- # ---------------------------------------------------------------------------
- # 3. Chart builder (Sunburst only)
- # ---------------------------------------------------------------------------
- def make_sunburst(df, title):
- fig = px.sunburst(
- df,
- path=["Dataset", "Species", "NodeType"],
- values="LogCount",
- branchvalues="total",
- custom_data=["Count"],
- title=title,
- )
- fig = apply_node_colors(fig)
- fig.update_traces(
- texttemplate="%{label}<br>%{customdata[0]:,}",
- hovertemplate="%{label}<br>Count: %{customdata[0]:,}<extra></extra>",
- insidetextorientation="radial",
- marker=dict(line=dict(color="white", width=1.5)),
- )
- fig.update_layout(
- font=dict(family="Arial, sans-serif", size=15),
- title=dict(x=0.5, font=dict(size=22, family="Arial, sans-serif")),
- margin=dict(t=80, l=10, r=10, b=10),
- paper_bgcolor="white",
- )
- return fig
- # ---------------------------------------------------------------------------
- # 4. Run for both schemes, sunburst only -- SVG output only
- # ---------------------------------------------------------------------------
- if __name__ == "__main__":
- for scheme_key, scheme_label in SCHEME_TITLE.items():
- df = build_hierarchy(scheme_key)
- sun_fig = make_sunburst(df, f"Knowledge-graph node types")
- try:
- sun_fig.write_image(f"sunburst_aging_biomedical_{scheme_key}.svg", width=1400, height=1400, scale=2)
- print(f"Saved: sunburst_aging_biomedical_{scheme_key}.svg")
- except Exception as e:
- print("SVG saved")
- sun_fig.show()
- # %%
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- import matplotlib.patches as mpatches
- import matplotlib.patheffects as pe
- # Keep SVG text as real, editable <text> elements instead of converting every
- # label/title/legend entry into vector path outlines (matplotlib's default,
- # "svg.fonttype" = "path", bakes text into shapes that can't be edited or
- # re-fonted in Illustrator/Inkscape). "none" keeps it as selectable/editable text.
- plt.rcParams["svg.fonttype"] = "none"
- # -----------------------------------------------------------------------------
- # 1. LOAD & PREP
- # -----------------------------------------------------------------------------
- # Aging: long-format file, already has an edge_type + dataset column
- # (same source file as before, just keep the "Aging" rows this time).
- aging_path = "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/Edgetype.csv"
- # Biomedical: wide-format file (NodeType/Relation x Species columns), needs
- # reshaping to long + an edge_type derived from the Relation name.
- biomedical_path = "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/Species_Biomedical_121_12M_RelationType_unique_counts.csv"
- EDGE_TYPES_IN_NAME = [
- "NegativelyAssociatedWith",
- "PositivelyAssociatedWith",
- "NotAssociatedWith",
- "NoEffect",
- "Promotes",
- "Inhibits",
- ]
- def infer_edge_type(relation):
- for et in EDGE_TYPES_IN_NAME:
- if et in relation:
- return et
- return "relation"
- # --- Aging (long format already) ---
- df_aging = pd.read_csv(aging_path)
- df_aging = df_aging[df_aging["dataset"] == "Aging"].copy()
- df_aging = df_aging[df_aging["edge_type"] != "AssociatedWith"]
- df_aging = df_aging[["Relation", "Species", "count", "edge_type"]]
- df_aging["dataset"] = "Aging"
- # --- Biomedical (wide format -> long) ---
- bio_wide = pd.read_csv(biomedical_path)
- bio_wide = bio_wide[~bio_wide["Relation"].isin(["Species_AssociatedWith_Nodes", "Total Triples"])].copy()
- species_cols = [c for c in bio_wide.columns if c != "Relation"]
- df_biomedical = bio_wide.melt(id_vars="Relation", value_vars=species_cols,
- var_name="Species", value_name="count")
- df_biomedical = df_biomedical[df_biomedical["count"] > 0].copy()
- df_biomedical["edge_type"] = df_biomedical["Relation"].apply(infer_edge_type)
- df_biomedical["dataset"] = "Biomedical"
- # --- combine ---
- df = pd.concat([df_aging, df_biomedical], ignore_index=True)
- # True log10(count) needs count > 0 (log10(0) is undefined) -- drop any
- # leftover zero/negative-count rows before taking logs anywhere below.
- df = df[df["count"] > 0].copy()
- # IMPORTANT: several raw Relation rows collapse into the same wedge (e.g. many
- # distinct relation names all map to edge_type "relation"). Sum their raw
- # counts into ONE number per (dataset, Species, edge_type) wedge FIRST, then
- # take a single log10 of that sum. Taking log10 of each row separately and
- # summing those logs would instead compute log10(count_1 * count_2 * ...) --
- # the log of a PRODUCT, not the log of the total -- which makes wedges with
- # many small rows balloon past wedges with one genuinely large count. This
- # way, wedge size is proportional to log10(true total count) for that wedge.
- df = (df.groupby(["dataset", "Species", "edge_type"], sort=False)["count"]
- .sum().reset_index())
- def format_count(n):
- if n >= 1_000_000_000:
- return f"{n/1_000_000_000:.2f}B"
- elif n >= 1_000_000:
- return f"{n/1_000_000:.2f}M"
- elif n >= 1_000:
- return f"{n/1_000:.1f}k"
- return str(int(n))
- # -----------------------------------------------------------------------------
- # 2. COLOR MAPS
- # -----------------------------------------------------------------------------
- # NOTE: Biomedical keeps the SAME teal previously used for EvoAge.
- dataset_colors = {"Aging": "#f4c2c2", "Biomedical": "#7fb3b8"}
- species_colors = {sp: "#e8e8e8" for sp in
- ["Human", "Yeast", "Celegans", "Drosophila", "Zebrafish", "Mouse"]}
- edge_colors = {
- "NoEffect": "#a8546b",
- "relation": "#e8ddc7",
- "Promotes": "#8ecae6",
- "PositivelyAssociatedWith": "#a3c98f",
- "NotAssociatedWith": "#b0b0b0",
- "NegativelyAssociatedWith": "#b39ddb",
- "Inhibits": "#e8a48c",
- # AssociatedWith intentionally excluded
- }
- edge_legend_labels = {
- "NoEffect": "no-effect",
- "relation": "relation",
- "Promotes": "promotes",
- "PositivelyAssociatedWith": "positively-associated",
- "NotAssociatedWith": "not-associated",
- "NegativelyAssociatedWith": "negatively-associated",
- "Inhibits": "inhibits",
- }
- # -----------------------------------------------------------------------------
- # 3. ANGLE ASSIGNMENT (recursive, log10 of the TRUE raw total at every level)
- # At each level (dataset -> Species -> edge_type), a node's weight is
- # log10(sum of its own raw counts) -- recomputed fresh from "count" every
- # time, not inherited by summing children's already-logged weights. That
- # means a dataset with a few huge counts (e.g. Biomedical, billions) will
- # correctly get a bigger wedge than one with many small/medium counts
- # (e.g. Aging), because we're logging the real total once, not summing
- # several small logs together (which silently favors "many small values").
- # -----------------------------------------------------------------------------
- def assign_angles(data, group_cols, start_angle=0, end_angle=360):
- if not group_cols:
- return data
- col = group_cols[0]
- # Weight of each child at THIS level = log10(sum of raw counts under it).
- # Recomputing from raw "count" at every level (rather than summing
- # already-logged child weights) is what keeps parent wedges honestly
- # proportional to what's actually inside them -- a dataset with a few
- # huge counts will correctly out-weigh one with many small counts.
- sums = data.groupby(col, sort=False)["count"].sum()
- weights = np.log10(sums)
- total_weight = weights.sum()
- order = weights.sort_values(ascending=False).index
- angle = start_angle
- frames = []
- for key in order:
- w = weights[key]
- span = (w / total_weight) * (end_angle - start_angle) if total_weight > 0 else 0
- sub = data[data[col] == key].copy()
- sub = assign_angles(sub, group_cols[1:], angle, angle + span)
- sub[f"{col}_start"] = angle
- sub[f"{col}_end"] = angle + span
- frames.append(sub)
- angle += span
- return pd.concat(frames)
- # -----------------------------------------------------------------------------
- # 4. DRAWING HELPERS
- # -----------------------------------------------------------------------------
- r_hole = 0.28
- r_dataset = (r_hole, 0.55)
- r_species = (0.58, 0.78)
- r_edge = (0.81, 1.05)
- r_label = 1.09
- def draw_wedge(ax, t1, t2, r_inner, r_outer, color, edgecolor="white", lw=0.8):
- if t2 - t1 <= 0:
- return
- w = mpatches.Wedge((0, 0), r_outer, t1, t2, width=r_outer - r_inner,
- facecolor=color, edgecolor=edgecolor, linewidth=lw)
- ax.add_patch(w)
- def draw_sunburst(ax, df, title):
- # ── Ring 1: dataset ──────────────────────────────────────────────────────
- for dataset, grp in df.groupby("dataset", sort=False):
- t1 = grp["dataset_start"].iloc[0]
- t2 = grp["dataset_end"].iloc[0]
- draw_wedge(ax, t1, t2, r_dataset[0], r_dataset[1],
- dataset_colors.get(dataset, "#cccccc"))
- mid = np.deg2rad((t1 + t2) / 2)
- r_t = (r_dataset[0] + r_dataset[1]) / 2
- ax.text(r_t * np.cos(mid), r_t * np.sin(mid), dataset,
- ha="center", va="center", fontsize=13, fontweight="bold")
- # ── Ring 2: species ───────────────────────────────────────────────────────
- for (dataset, sp), grp in df.groupby(["dataset", "Species"], sort=False):
- t1 = grp["Species_start"].iloc[0]
- t2 = grp["Species_end"].iloc[0]
- draw_wedge(ax, t1, t2, r_species[0], r_species[1],
- species_colors.get(sp, "#e8e8e8"))
- mid_deg = (t1 + t2) / 2
- mid = np.deg2rad(mid_deg)
- r_t = (r_species[0] + r_species[1]) / 2
- norm = mid_deg % 360
- rot = mid_deg + 180 if 90 < norm < 270 else mid_deg
- ax.text(r_t * np.cos(mid), r_t * np.sin(mid), sp,
- ha="center", va="center", fontsize=9,
- rotation=rot, rotation_mode="anchor")
- # ── Ring 3: edge_type ─────────────────────────────────────────────────────
- for _, row in df.iterrows():
- t1, t2 = row["angle_start"], row["angle_end"]
- color = edge_colors.get(row["edge_type"], "#cccccc")
- draw_wedge(ax, t1, t2, r_edge[0], r_edge[1], color)
- # ── Outer count labels ────────────────────────────────────────────────────
- # IMPORTANT: multiple raw rows can share the same edge_type wedge (e.g. many
- # different relation sub-types all rolled into the single "relation" wedge),
- # and they all share the same angle_start/angle_end. Looping over raw rows
- # here would draw one label per underlying row, all stacked at the same
- # midpoint -> garbled overlapping digits. Aggregate to one label per wedge first.
- label_df = (df.groupby(["Species", "edge_type", "angle_start", "angle_end"], sort=False)
- ["count"].sum().reset_index())
- MIN_LABEL_DEG = 1.5
- skipped = []
- for _, row in label_df.iterrows():
- t1, t2 = row["angle_start"], row["angle_end"]
- if t2 - t1 <= 0:
- continue
- if (t2 - t1) < MIN_LABEL_DEG:
- skipped.append((row.get("Species", ""), row["edge_type"], row["count"]))
- continue
- mid_deg = (t1 + t2) / 2
- mid = np.deg2rad(mid_deg)
- norm = mid_deg % 360
- rot = mid_deg + 180 if 90 < norm < 270 else mid_deg
- ha = "right" if 90 < norm < 270 else "left"
- ax.text(r_label * np.cos(mid), r_label * np.sin(mid),
- format_count(row["count"]),
- ha=ha, va="center", fontsize=9, fontweight="bold", color="black",
- rotation=rot, rotation_mode="anchor",
- path_effects=[pe.withStroke(linewidth=2.5, foreground="white")])
- # list skipped (too-thin-to-label) values in the corner so the exact
- # numbers are still available, just not crammed onto the wedge
- if skipped:
- lines = [f"{sp} / {et}: {format_count(c)}" for sp, et, c in skipped]
- ax.text(-1.55, -1.55, "small slices (not labeled on ring):\n" + "\n".join(lines),
- ha="left", va="bottom", fontsize=6, color="dimgray")
- ax.set_xlim(-1.6, 1.6)
- ax.set_ylim(-1.6, 1.6)
- ax.axis("off")
- ax.set_title(title, fontsize=15, fontweight="bold", pad=20)
- # -----------------------------------------------------------------------------
- # 5. BUILD LOG-SCALED DATAFRAME (only)
- # -----------------------------------------------------------------------------
- df_log = assign_angles(df.copy(), ["dataset", "Species", "edge_type"])
- df_log["angle_start"] = df_log["edge_type_start"]
- df_log["angle_end"] = df_log["edge_type_end"]
- # -----------------------------------------------------------------------------
- # 6. PLOT — single log-scaled sunburst
- # -----------------------------------------------------------------------------
- fig, ax = plt.subplots(1, 1, figsize=(13, 13), subplot_kw={"aspect": "equal"})
- draw_sunburst(ax, df_log,
- "Edge-Type Counts — Log₁₀-scaled\n(wedge size ∝ log₁₀(count))")
- legend_handles = [
- mpatches.Patch(color=c, label=edge_legend_labels.get(k, k))
- for k, c in edge_colors.items()
- ]
- legend_handles += [
- mpatches.Patch(color=c, label=k)
- for k, c in dataset_colors.items()
- ]
- fig.legend(handles=legend_handles, loc="lower center", ncol=5,
- fontsize=11, bbox_to_anchor=(0.5, -0.03), frameon=False)
- fig.suptitle("Sunburst: Edge-Type Counts by Dataset & Species\n(AssociatedWith excluded)",
- fontsize=17, fontweight="bold", y=1.01)
- plt.tight_layout()
- plt.savefig("edge-type-sunburst_log_aging_biomedical.svg", format="svg", bbox_inches="tight")
- plt.show()
- df.head()
- # %% [markdown]
- # ## 1e
- # %%
- import matplotlib
- matplotlib.rcParams['svg.fonttype'] = 'none'
- matplotlib.rcParams['font.family'] = 'sans-serif'
- import matplotlib.pyplot as plt
- import matplotlib.patches as mpatches
- import matplotlib.lines as mlines
- import numpy as np
- import pandas as pd
- from matplotlib.patches import Circle, Ellipse
- # ── File paths ────────────────────────────────────────────────────────────────
- AGING_NODES_CSV = "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/aging_121_12M_Species_NodeType_unique_counts.csv"
- AGING_TRIPLES_CSV = "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/aging_121_12M_RelationType_AllSpecies.csv"
- BIO_NODES_CSV = "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/Species_Biomedical_121_12M_NodeType_unique_counts.csv"
- BIO_TRIPLES_CSV = "/storage/Arushi/090526_EvoAge/kg_formation/species_wise_kg_building/Species_Biomedical_121_12M_RelationType_unique_counts.csv"
- OUTPUT_SVG = "nodes_triples_dumbbell.svg"
- OUTPUT_PNG = "nodes_triples_dumbbell.png"
- # ── Colors ────────────────────────────────────────────────────────────────────
- AGING_COL = '#F08080' # salmon
- BIO_COL = '#80b4b9' # teal
- # ── Species order (left → right on x-axis, matching icon order) ───────────────
- # Column names as they appear in CSVs → display label
- SPECIES_MAP = {
- 'Human': 'Human',
- 'Mouse': 'Mouse',
- 'Zebrafish': 'Zebrafish',
- 'Celegans': 'C. elegans',
- 'Drosophila': 'Drosophila',
- 'Yeast': 'Yeast',
- }
- SPECIES_COLS = list(SPECIES_MAP.keys()) # CSV column names
- SPECIES_ORDER = list(SPECIES_MAP.keys()) # plot order
- # ── Load & extract TOTAL row ──────────────────────────────────────────────────
- def load_total_row(filepath, total_col_value):
- """
- Reads a CSV where the first column is a label column.
- Returns a dict {species_col: count} for the row whose
- first-column value matches `total_col_value`.
- """
- df = pd.read_csv(filepath)
- label_col = df.columns[0]
- row = df[df[label_col] == total_col_value].iloc[0]
- return {sp: int(row[sp]) if sp in df.columns else 0 for sp in SPECIES_COLS}
- aging_nodes = load_total_row(AGING_NODES_CSV, 'TOTAL')
- aging_triples = load_total_row(AGING_TRIPLES_CSV, 'Total Triples')
- bio_nodes = load_total_row(BIO_NODES_CSV, 'TOTAL')
- bio_triples = load_total_row(BIO_TRIPLES_CSV, 'Total Triples')
- # ── log10(count + 1) helper ───────────────────────────────────────────────────
- def log1(val):
- return np.log10(val + 1)
- # ── Figure ────────────────────────────────────────────────────────────────────
- x_pos = np.arange(len(SPECIES_ORDER), dtype=float)
- OFFSET = 0.15 # left stem = nodes, right stem = triples
- fig, ax = plt.subplots(figsize=(7.5, 6.2))
- fig.patch.set_facecolor('white')
- ax.set_facecolor('white')
- for i, sp in enumerate(SPECIES_ORDER):
- cx = x_pos[i]
- an = log1(aging_nodes[sp])
- bn = log1(bio_nodes[sp])
- at = log1(aging_triples[sp])
- bt = log1(bio_triples[sp])
- x_node = cx - OFFSET # left dumbbell → node counts
- x_triple = cx + OFFSET # right dumbbell → triple counts
- # vertical stems
- ax.plot([x_node, x_node], [an, bn], color='#333333', lw=1.3, zorder=1, solid_capstyle='round')
- ax.plot([x_triple, x_triple], [at, bt], color='#333333', lw=1.3, zorder=1, solid_capstyle='round')
- # node dumbbell: circles
- ax.scatter(x_node, an, marker='o', s=90, color=AGING_COL, zorder=4, linewidths=0.5, edgecolors='white')
- ax.scatter(x_node, bn, marker='o', s=90, color=BIO_COL, zorder=4, linewidths=0.5, edgecolors='white')
- # triple dumbbell: diamonds
- ax.scatter(x_triple, at, marker='D', s=70, color=AGING_COL, zorder=4, linewidths=0.5, edgecolors='white')
- ax.scatter(x_triple, bt, marker='D', s=70, color=BIO_COL, zorder=4, linewidths=0.5, edgecolors='white')
- # ── Axes formatting ───────────────────────────────────────────────────────────
- ax.set_ylabel(r'$\log_{10}\ (\mathrm{count}+1)$', fontsize=12)
- ax.set_xticks(x_pos)
- # Add species names as x-axis labels
- ax.set_xticklabels([SPECIES_MAP[sp] for sp in SPECIES_ORDER], fontsize=10, rotation=45, ha='right')
- ax.set_xlim(-0.6, len(SPECIES_ORDER) - 0.4)
- ax.set_ylim(-1.7, 10.5)
- ax.set_yticks(range(0, 11, 2))
- ax.tick_params(axis='y', labelsize=11)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- ax.spines['bottom'].set_linewidth(1.2)
- ax.spines['left'].set_linewidth(1.2)
- ax.set_title('Distribution (Nodes, Triples)', fontsize=13, pad=10)
- # ── Legend ────────────────────────────────────────────────────────────────────
- aging_patch = mpatches.Patch(color=AGING_COL, label='Aging')
- bio_patch = mpatches.Patch(color=BIO_COL, label='Biomedical')
- node_marker = mlines.Line2D([], [], color='grey', marker='o', linestyle='None', markersize=8, label='Node count')
- tri_marker = mlines.Line2D([], [], color='grey', marker='D', linestyle='None', markersize=7, label='Triple count')
- ax.legend(
- handles=[aging_patch, node_marker, bio_patch, tri_marker],
- loc='upper right', fontsize=10, frameon=False, ncol=2,
- handlelength=1.2, columnspacing=0.8, labelspacing=0.4,
- )
- # ── Save ──────────────────────────────────────────────────────────────────────
- plt.tight_layout()
- plt.savefig(OUTPUT_SVG, format='svg', bbox_inches='tight')
- plt.savefig(OUTPUT_PNG, dpi=200, bbox_inches='tight')
- print(f"Saved: {OUTPUT_SVG}, {OUTPUT_PNG}")
- # %% [markdown]
- # ## 1f
- # %%
- # /storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Aging_121_12M/Aging_121_12M_degree.csv
- # /storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/EvoAge_121_12M/EvoAge_121_12M_degree.csv
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- from matplotlib.ticker import LogLocator, FuncFormatter
- # -----------------------------------------------------------------------------
- # 1. LOAD DATA
- # -----------------------------------------------------------------------------
- aging_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Aging_121_12M/Aging_121_12M_degree.csv"
- evoage_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/EvoAge_121_12M/EvoAge_121_12M_degree.csv"
- aging = pd.read_csv(aging_path)
- evoage = pd.read_csv(evoage_path)
- # -----------------------------------------------------------------------------
- # 2. BUILD NORMALIZED DEGREE DISTRIBUTION: P(k) = fraction of nodes with degree k
- # -----------------------------------------------------------------------------
- def to_distribution(df, model_name):
- N = len(df)
- dist = (
- df["total_degree"]
- .value_counts()
- .rename_axis("degree")
- .reset_index(name="num_nodes")
- )
- dist["P_k"] = dist["num_nodes"] / N # normalize by total node count
- dist["Model"] = model_name
- dist["N"] = N
- return dist
- aging_dist = to_distribution(aging, "Aging")
- evoage_dist = to_distribution(evoage, "EvoAge")
- N_aging = len(aging)
- N_evoage = len(evoage)
- data_long = pd.concat([evoage_dist, aging_dist], ignore_index=True)
- data_long = data_long[(data_long["degree"] > 0) & (data_long["P_k"] > 0)]
- # -----------------------------------------------------------------------------
- # 3. PLOT (log-log scatter)
- # -----------------------------------------------------------------------------
- colors = {"EvoAge": "#2c7189", "Aging": "#b26f77"}
- labels = {"EvoAge": f"EvoAge (N = {N_evoage:,})",
- "Aging": f"Aging (N = {N_aging:,})"}
- fig, ax = plt.subplots(figsize=(10, 8))
- for model in ["EvoAge", "Aging"]:
- subset = data_long[data_long["Model"] == model]
- ax.scatter(
- subset["degree"], subset["P_k"],
- color=colors[model], s=45, alpha=0.7,
- edgecolor="none", label=labels[model],
- )
- ax.set_xscale("log")
- ax.set_yscale("log")
- ax.xaxis.set_major_locator(LogLocator(base=10, numticks=15))
- ax.yaxis.set_major_locator(LogLocator(base=10, numticks=15))
- ax.xaxis.set_major_formatter(FuncFormatter(lambda x, _: f"$10^{{{int(np.log10(x))}}}$"))
- ax.yaxis.set_major_formatter(FuncFormatter(lambda x, _: f"$10^{{{int(np.log10(x))}}}$"))
- # Auto-fit limits to actual data
- max_deg = data_long["degree"].max()
- ax.set_xlim(1, 10 ** np.ceil(np.log10(max_deg)))
- min_pk = data_long["P_k"].min()
- ax.set_ylim(10 ** np.floor(np.log10(min_pk)), 1)
- ax.set_title("Comparative Log-Log Plot of Degree Distributions",
- fontsize=20, fontweight="bold", pad=15)
- ax.text(0.5, 1.02, "Normalized scatter representation on log-log scale",
- transform=ax.transAxes, ha="center", fontsize=12, color="gray")
- ax.set_xlabel("Node Degree, $k$ (Log Scale)", fontsize=16)
- ax.set_ylabel("Fraction of Nodes, $P(k)$ (Log Scale)", fontsize=16)
- ax.spines["top"].set_visible(False)
- ax.spines["right"].set_visible(False)
- ax.spines["left"].set_color("black")
- ax.spines["bottom"].set_color("black")
- ax.tick_params(axis="both", which="major", length=6, width=1, color="black", labelsize=13)
- ax.tick_params(axis="both", which="minor", length=0)
- ax.grid(False, which="major")
- ax.grid(False, which="minor")
- ax.legend(loc="lower center", bbox_to_anchor=(0.5, -0.18),
- ncol=2, frameon=False, fontsize=14, markerscale=1.5)
- plt.tight_layout()
- # -----------------------------------------------------------------------------
- # 4. SAVE AS PDF
- # -----------------------------------------------------------------------------
- plt.savefig("FIG1/Fig1_h_main.pdf", format="pdf", bbox_inches="tight")
- plt.savefig("FIG1/fig1h-degree.png", format="png", bbox_inches="tight")
- plt.show()
- # %% [markdown]
- # ## 1g
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- # -----------------------------------------------------------------------------
- # 1. LOAD DATA
- # -----------------------------------------------------------------------------
- aging_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Aging_121_12M/Aging_121_12M_louvain_partitions.csv"
- evoage_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Biomedical_121_12M/Biomedical_121_12M_louvain_partitions.csv"
- def read_louvain_csv(path):
- df = pd.read_csv(path)
- df.columns = df.columns.str.lower()
- if not {"vertex", "partition"}.issubset(df.columns):
- raise ValueError(f"Expected 'vertex' and 'partition' columns in {path}, got {list(df.columns)}")
- return df[["vertex", "partition"]]
- aging_raw = read_louvain_csv(aging_path)
- evoage_raw = read_louvain_csv(evoage_path)
- # -----------------------------------------------------------------------------
- # 2. COMMUNITY SIZE TABLES
- # -----------------------------------------------------------------------------
- def community_sizes(df):
- return (
- df.groupby("partition")
- .size()
- .reset_index(name="size")
- .sort_values("size", ascending=False)
- .reset_index(drop=True)
- )
- aging_sizes = community_sizes(aging_raw)
- evoage_sizes = community_sizes(evoage_raw)
- N_aging = len(aging_raw)
- N_evoage = len(evoage_raw)
- # -----------------------------------------------------------------------------
- # 3. QUICK STATS (mirrors R's cat block)
- # -----------------------------------------------------------------------------
- def entropy(sizes):
- p = sizes / sizes.sum()
- p = p[p > 0]
- return float(-np.sum(p * np.log(p)))
- print("\n=== Summary ===")
- print(f"Aging | Nodes: {N_aging:,} | Communities: {len(aging_sizes):,} | "
- f"Largest: {100*aging_sizes['size'].max()/N_aging:.2f}% | "
- f"Entropy: {entropy(aging_sizes['size'].values):.3f}")
- print(f"Biomedical | Nodes: {N_evoage:,} | Communities: {len(evoage_sizes):,} | "
- f"Largest: {100*evoage_sizes['size'].max()/N_evoage:.2f}% | "
- f"Entropy: {entropy(evoage_sizes['size'].values):.3f}\n")
- # -----------------------------------------------------------------------------
- # 4. TOP-K COMMUNITIES
- # -----------------------------------------------------------------------------
- K = 15
- top_aging = aging_sizes.head(K).copy()
- top_aging["rank"] = np.arange(1, len(top_aging) + 1)
- top_aging["pct"] = 100 * top_aging["size"] / N_aging
- top_aging["graph"] = "Aging"
- top_evoage = evoage_sizes.head(K).copy()
- top_evoage["rank"] = np.arange(1, len(top_evoage) + 1)
- top_evoage["pct"] = 100 * top_evoage["size"] / N_evoage
- top_evoage["graph"] = "Biomedical"
- top_df = pd.concat([top_aging, top_evoage], ignore_index=True)
- # -----------------------------------------------------------------------------
- # 5. PLOT — grouped bar chart
- # -----------------------------------------------------------------------------
- colors = {"Aging": "#eeb4b4", "Biomedical": "#7ebcdb"}
- fig, ax = plt.subplots(figsize=(10, 6))
- ranks = np.arange(1, K + 1)
- width = 0.4
- aging_vals = top_df[top_df["graph"] == "Aging"].set_index("rank").reindex(ranks)["pct"].fillna(0)
- evoage_vals = top_df[top_df["graph"] == "Biomedical"].set_index("rank").reindex(ranks)["pct"].fillna(0)
- ax.bar(ranks - width/2, aging_vals, width=width, color=colors["Aging"], label="Aging",
- edgecolor="none")
- ax.bar(ranks + width/2, evoage_vals, width=width, color=colors["Biomedical"], label="Biomedical",
- edgecolor="none")
- ax.set_xticks(ranks)
- ax.set_xticklabels(ranks, fontsize=11)
- ax.set_title(f"Top-{K} Community Sizes (as % of nodes)",
- fontsize=15, fontweight="bold", pad=10)
- ax.text(0.5, 1.02,
- f"Top-{K} communities by size, shown as % of total nodes. "
- "Larger bars indicate more concentration in a few communities.",
- transform=ax.transAxes, ha="center", fontsize=10, color="gray")
- ax.set_xlabel("Community rank (by size, 1 = largest)", fontsize=13)
- ax.set_ylabel("Share of nodes in community (%)", fontsize=13)
- # theme_classic-ish
- ax.spines["top"].set_visible(False)
- ax.spines["right"].set_visible(False)
- ax.spines["left"].set_linewidth(0.7)
- ax.spines["bottom"].set_linewidth(0.7)
- ax.tick_params(axis="both", which="major", length=5, width=0.7, color="black", labelsize=11)
- ax.grid(False)
- ax.legend(loc="upper right", frameon=False, fontsize=12, ncol=2)
- plt.tight_layout()
- # -----------------------------------------------------------------------------
- # 6. SAVE AS PDF
- # -----------------------------------------------------------------------------
- plt.savefig("FIG1/aging-biomedical-commmunity.pdf", format="pdf", bbox_inches="tight")
- plt.show()
- # %% [markdown]
- # ## 1h
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- from matplotlib.ticker import LogLocator, FuncFormatter
- # -----------------------------------------------------------------------------
- # 1. LOAD DATA
- # -----------------------------------------------------------------------------
- aging_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Aging_121_12M/Aging_121_12M_betweenness_centrality_sorted.csv"
- evoage_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Biomedical_121_12M/Biomedical_121_12M_betweenness_centrality_sorted.csv"
- df_aging = pd.read_csv(aging_path)
- df_evoage = pd.read_csv(evoage_path)
- # -----------------------------------------------------------------------------
- # 2. CONFIG
- # -----------------------------------------------------------------------------
- TARGET_ROWS = 2_000_000
- NBINS_PDF = 120
- SEED = 42
- # -----------------------------------------------------------------------------
- # 3. HELPERS
- # -----------------------------------------------------------------------------
- def positive_vec(x, min_positive=1e-12):
- """Keep only finite, positive values (required for log-scale)."""
- x = pd.to_numeric(x, errors="coerce").to_numpy()
- x = x[np.isfinite(x)]
- x = x[x > 0]
- return x if len(x) else np.array([min_positive])
- def log_binned_pdf(arr, nbins=120):
- """Width-corrected PDF using log-spaced bins."""
- lo, hi = arr.min(), arr.max()
- if hi <= lo:
- return np.array([hi]), np.array([1.0])
- bins = np.logspace(np.log10(lo), np.log10(hi), nbins + 1)
- counts, edges = np.histogram(arr, bins=bins)
- widths = np.diff(edges)
- total = counts.sum()
- pdf = np.where(widths > 0, (counts / total) / widths, 0)
- centers = np.sqrt(edges[:-1] * edges[1:]) # geometric mean of bin edges
- # keep only bins with non-zero density (log axis can't plot 0)
- mask = pdf > 0
- return centers[mask], pdf[mask]
- # -----------------------------------------------------------------------------
- # 4. PREP
- # -----------------------------------------------------------------------------
- # Subsample EvoAge if too large (matches R's TARGET_ROWS = 2e6 on df2)
- if len(df_evoage) > TARGET_ROWS:
- df_evoage = df_evoage.sample(n=TARGET_ROWS, random_state=SEED)
- a_aging = positive_vec(df_aging["betweenness_centrality"])
- a_evoage = positive_vec(df_evoage["betweenness_centrality"])
- x_aging, y_aging = log_binned_pdf(a_aging, NBINS_PDF)
- x_evoage, y_evoage = log_binned_pdf(a_evoage, NBINS_PDF)
- # -----------------------------------------------------------------------------
- # 5. PLOT
- # -----------------------------------------------------------------------------
- colors = {"Aging": "#eeb4b4", "Biomedical": "#7ebcdb"}
- fig, ax = plt.subplots(figsize=(8, 7))
- ax.plot(x_evoage, y_evoage, color=colors["Biomedical"], linewidth=1.4, label="Biomedical")
- ax.plot(x_aging, y_aging, color=colors["Aging"], linewidth=1.4, label="Aging")
- ax.set_xscale("log")
- ax.set_yscale("log")
- # 10^n tick formatting
- def sci_fmt(x, _):
- return f"$10^{{{int(np.log10(x))}}}$"
- ax.xaxis.set_major_locator(LogLocator(base=10, numticks=20))
- ax.yaxis.set_major_locator(LogLocator(base=10, numticks=20))
- ax.xaxis.set_major_formatter(FuncFormatter(sci_fmt))
- ax.yaxis.set_major_formatter(FuncFormatter(sci_fmt))
- # Labels
- ax.set_title("connectivity", fontsize=16, pad=12)
- ax.set_xlabel("betweenness centrality", fontsize=14)
- ax.set_ylabel("density", fontsize=14)
- # Rotate x tick labels a bit (matches your reference image)
- plt.setp(ax.get_xticklabels(), rotation=45, ha="right")
- # Clean minimal look
- ax.spines["top"].set_visible(False)
- ax.spines["right"].set_visible(False)
- ax.tick_params(axis="both", which="major", length=5, width=1, color="black", labelsize=11)
- ax.tick_params(axis="both", which="minor", length=0)
- ax.grid(False)
- # Legend top-center inside plot
- ax.legend(loc="upper center", bbox_to_anchor=(0.5, 0.98),
- ncol=2, frameon=False, fontsize=13)
- plt.tight_layout()
- # -----------------------------------------------------------------------------
- # 6. SAVE AS PDF
- # -----------------------------------------------------------------------------
- plt.savefig("FIG1/aging-biomedical-centrality.pdf", format="pdf", bbox_inches="tight")
- plt.show()
- # %% [markdown]
- # ## 1i
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- from matplotlib.ticker import LogLocator, FuncFormatter
- # -----------------------------------------------------------------------------
- # 1. LOAD DATA
- # -----------------------------------------------------------------------------
- aging_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Aging_121_12M_vertex_counts.csv"
- evoage_path = "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/Graph_features/Output_csv/Biomedical_121_12M/Biomedical_121_12M_vertex_counts.csv"
- def read_tri_csv(path):
- df = pd.read_csv(path)
- df.columns = df.columns.str.lower()
- if not {"vertex", "counts"}.issubset(df.columns):
- raise ValueError(f"Expected 'vertex' and 'counts' columns in {path}, got {list(df.columns)}")
- return df[["vertex", "counts"]].apply(pd.to_numeric, errors="coerce")
- aging_raw = read_tri_csv(aging_path)
- evoage_raw = read_tri_csv(evoage_path)
- N_aging = len(aging_raw)
- N_evoage = len(evoage_raw)
- # -----------------------------------------------------------------------------
- # 2. QUICK STATS
- # -----------------------------------------------------------------------------
- aging_zero = (aging_raw["counts"] <= 0).mean()
- evoage_zero = (evoage_raw["counts"] <= 0).mean()
- print("\n=== Triangle Count Summary ===")
- print(f"Aging | Nodes: {N_aging:,} | zero-triangle nodes: {100*aging_zero:.2f}%")
- print(f"Biomedical | Nodes: {N_evoage:,} | zero-triangle nodes: {100*evoage_zero:.2f}%\n")
- # Keep positive counts for log-scale
- pos_aging = aging_raw.loc[aging_raw["counts"] > 0, "counts"].to_numpy()
- pos_evoage = evoage_raw.loc[evoage_raw["counts"] > 0, "counts"].to_numpy()
- # -----------------------------------------------------------------------------
- # 3. LOG-BINNED PDF (matches R's hist(..., include.lowest=TRUE, right=FALSE))
- # -----------------------------------------------------------------------------
- def log_binned_pdf(arr, nbins=120):
- arr = arr[np.isfinite(arr) & (arr > 0)]
- if len(arr) == 0:
- return np.array([1.0]), np.array([1.0])
- lo, hi = arr.min(), arr.max()
- if hi <= lo:
- return np.array([hi]), np.array([1.0])
- bins = np.logspace(np.log10(lo), np.log10(hi), nbins + 1)
- # Replicate R's right=FALSE, include.lowest=TRUE:
- # left-closed, right-open intervals [a,b); include max value in last bin
- idx = np.digitize(arr, bins, right=False)
- idx[arr == bins[-1]] = nbins # include the maximum value in the last bin
- counts = np.bincount(idx, minlength=nbins + 2)[1:nbins + 1]
- widths = np.diff(bins)
- total = counts.sum()
- pdf = np.where(widths > 0, (counts / total) / widths, 0)
- centers = np.sqrt(bins[:-1] * bins[1:])
- mask = pdf > 0
- return centers[mask], pdf[mask]
- x_aging, y_aging = log_binned_pdf(pos_aging, nbins=120)
- x_evoage, y_evoage = log_binned_pdf(pos_evoage, nbins=120)
- # -----------------------------------------------------------------------------
- # 4. PLOT
- # -----------------------------------------------------------------------------
- colors = {"Aging": "#eeb4b4", "Biomedical": "#7ebcdb"}
- fig, ax = plt.subplots(figsize=(8, 7))
- # Aging as dashed line (matches your reference image)
- ax.plot(x_aging, y_aging, color=colors["Aging"],
- linewidth=1.4, linestyle="--", label="Aging")
- ax.plot(x_evoage, y_evoage, color=colors["Biomedical"],
- linewidth=1.4, linestyle="-", label="Biomedical")
- ax.set_xscale("log")
- ax.set_yscale("log")
- def sci_fmt(x, _):
- return f"$10^{{{int(np.log10(x))}}}$"
- ax.xaxis.set_major_locator(LogLocator(base=10, numticks=20))
- ax.yaxis.set_major_locator(LogLocator(base=10, numticks=20))
- ax.xaxis.set_major_formatter(FuncFormatter(sci_fmt))
- ax.yaxis.set_major_formatter(FuncFormatter(sci_fmt))
- ax.set_title("clustering", fontsize=16, pad=12)
- ax.set_xlabel("triangles per node", fontsize=14)
- ax.set_ylabel("density", fontsize=14)
- ax.spines["top"].set_visible(False)
- ax.spines["right"].set_visible(False)
- ax.tick_params(axis="both", which="major", length=5, width=1, color="black", labelsize=11)
- ax.tick_params(axis="both", which="minor", length=0)
- ax.grid(False)
- ax.legend(loc="upper center", bbox_to_anchor=(0.5, 0.98),
- ncol=2, frameon=False, fontsize=13)
- plt.tight_layout()
- plt.show()
- # -----------------------------------------------------------------------------
- # 5. SAVE AS PDF
- # -----------------------------------------------------------------------------
- plt.savefig("FIG1/aging-biomedical-clustering.pdf", format="pdf", bbox_inches="tight")
- # %% [markdown]
- # ## Main Figure 2
- # %% [markdown]
- # ## 2b
- # %%
- """
- 3-set Venn across both KG variants (1:1 vs 121_12M, sharing the same test set),
- for all three KG families — Aging, Biomedical, EvoAge.
- """
- import gc
- import numpy as np
- import torch
- import matplotlib.pyplot as plt
- from matplotlib_venn import venn3
- # ============================================================
- # CONFIG
- # ============================================================
- FAMILIES = {
- "Aging": {
- "full_1to1": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_aging_kg_new/Store_House/Aging_specific_1_to_1_KG.pt",
- "test": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_aging_kg_new/Store_House/Aging_specific_1to1_KG_test_10_shared.pt",
- "full_12m": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_aging_kg_new/Store_House/Aging_specific_121_12M_KG.pt",
- },
- "Biomedical": {
- "full_1to1": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_biomedical_kg_new/Store_House/Biomedical_1_to_1_KG.pt",
- "test": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_biomedical_kg_new/Store_House/Biomedical_1to1_KG_test_10_shared.pt",
- "full_12m": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_biomedical_kg_new/Store_House/Biomedical_121_12M_KG.pt",
- },
- "EvoAge": {
- "full_1to1": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_evoage_kg_new/Store_House/EvoAge_1_to_1_KG.pt",
- "test": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_evoage_kg_new/Store_House/EvoAge_1to1_KG_test_10_shared.pt",
- "full_12m": "/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_evoage_kg_new/Store_House/EvoAge_121_12M_to_many_KG.pt",
- },
- }
- ASSUME_UNIQUE_INPUT = False # True -> sort only, skip dedupe (if files already unique)
- UNWEIGHTED = True # True -> readable fixed layout; False -> area-proportional
- USE_GPU = False # True -> CuPy path (needs cupy + enough VRAM per set)
- xp = np
- if USE_GPU:
- import cupy as cp
- xp = cp
- # ============================================================
- # LOADER -> (N,3) int64 on CPU
- # ============================================================
- def load_triples_pt(path):
- obj = torch.load(path, map_location="cpu", weights_only=False)
- if isinstance(obj, torch.Tensor) and obj.ndim == 2 and obj.shape[1] == 3:
- return obj.numpy().astype(np.int64, copy=False)
- if isinstance(obj, dict):
- for key in ("triples", "edge_triples", "all_triples"):
- if key in obj:
- return np.asarray(obj[key], dtype=np.int64)
- if "edge_index" in obj and "edge_type" in obj:
- h, t = np.asarray(obj["edge_index"]); r = np.asarray(obj["edge_type"])
- return np.stack([h, r, t], axis=1).astype(np.int64)
- if hasattr(obj, "edge_index") and hasattr(obj, "edge_type"):
- h, t = obj.edge_index.numpy(); r = obj.edge_type.numpy()
- return np.stack([h, r, t], axis=1).astype(np.int64)
- raise ValueError(f"Unrecognized structure in {path}: type={type(obj)}")
- # ============================================================
- # helpers
- # ============================================================
- def unique_keys(triples, r_max, t_max):
- h = xp.asarray(triples[:, 0]); r = xp.asarray(triples[:, 1]); t = xp.asarray(triples[:, 2])
- keys = (h * r_max + r) * t_max + t # bijective packing
- return xp.sort(keys) if ASSUME_UNIQUE_INPUT else xp.unique(keys)
- def isin_mask(query, ref):
- """Which elements of sorted query are present in sorted ref."""
- if ref.size == 0:
- return xp.zeros(query.size, dtype=bool)
- idx = xp.clip(xp.searchsorted(ref, query), 0, ref.size - 1)
- return ref[idx] == query
- def count_in(query, ref): # pass the SMALLER array as query
- return int(isin_mask(query, ref).sum())
- # version-robust unweighted layout
- def make_venn3(subsets, labels, ax, unweighted=True):
- if not unweighted:
- return venn3(subsets=subsets, set_labels=labels, ax=ax)
- try:
- # newer matplotlib_venn (>=1.0): fixed_subset_sizes, no normalize_to
- from matplotlib_venn.layout.venn3 import DefaultLayoutAlgorithm
- layout = DefaultLayoutAlgorithm(fixed_subset_sizes=(1, 1, 1, 1, 1, 1, 1))
- return venn3(subsets=subsets, set_labels=labels, ax=ax, layout_algorithm=layout)
- except ImportError:
- # older matplotlib_venn: the deprecated helper still works
- from matplotlib_venn import venn3_unweighted
- return venn3_unweighted(subsets=subsets, set_labels=labels, ax=ax)
- # ============================================================
- # per-family Venn
- # ============================================================
- def venn_for_family(family, paths):
- raw = {
- "Full 1:1 KG": load_triples_pt(paths["full_1to1"]),
- "Test (shared)": load_triples_pt(paths["test"]),
- "Full 121_12M KG": load_triples_pt(paths["full_12m"]),
- }
- # shared packing bases within THIS family so keys are comparable across its three sets
- r_max = max(int(a[:, 1].max()) for a in raw.values()) + 1
- t_max = max(int(a[:, 2].max()) for a in raw.values()) + 1
- h_max = max(int(a[:, 0].max()) for a in raw.values())
- assert h_max * r_max * t_max + (r_max - 1) * t_max + (t_max - 1) < 2**63 - 1, \
- f"[{family}] packing overflows int64 — switch to a row-void view for keys"
- keysets = {}
- for name in list(raw):
- keysets[name] = unique_keys(raw[name], r_max, t_max)
- del raw[name]; gc.collect() # free each large raw array immediately
- A, B, C = keysets["Full 1:1 KG"], keysets["Test (shared)"], keysets["Full 121_12M KG"]
- a, b, c = A.size, B.size, C.size
- # intersections — always query the smaller array
- inBA = isin_mask(B, A); ab = int(inBA.sum())
- AB = B[inBA] # A∩B keys (small)
- bc = count_in(B, C)
- ac = count_in(A, C) if a <= c else count_in(C, A)
- abc = count_in(AB, C)
- # 7 disjoint regions
- ABC = abc
- AB_only = ab - abc
- AC_only = ac - abc
- BC_only = bc - abc
- A_only = a - ab - ac + abc
- B_only = b - ab - bc + abc
- C_only = c - ac - bc + abc
- print(f"\n=== {family} ===")
- for name, s in [("Full 1:1 KG", a), ("Test (shared)", b), ("Full 121_12M KG", c)]:
- print(f" {name:18s}: {s:,} unique triples")
- print(f" A∩B={ab:,} A∩C={ac:,} B∩C={bc:,} A∩B∩C={abc:,}")
- # matplotlib_venn subset order: (Abc, aBc, ABc, abC, AbC, aBC, ABC)
- subsets = (A_only, B_only, AB_only, C_only, AC_only, BC_only, ABC)
- labels = ("Full 1:1 KG", "Test (shared)", "Full 121_12M KG")
- fig, ax = plt.subplots(figsize=(9, 9))
- make_venn3(subsets, labels, ax, unweighted=UNWEIGHTED)
- ax.set_title(f"{family}: triple-set provenance across KG variants (real counts)",
- fontsize=15, fontweight="bold")
- fig.tight_layout()
- out = f"venn_3set_{family}.svg"
- fig.savefig(out, bbox_inches="tight")
- print(f" saved -> {out}")
- return fig
- # %%
- family = "Aging"
- fig_aging = venn_for_family(family, FAMILIES[family])
- plt.show()
- # %%
- family = "Biomedical"
- fig_biomedical = venn_for_family(family, FAMILIES[family])
- plt.show()
- # %%
- family = "EvoAge"
- fig_evoage = venn_for_family(family, FAMILIES[family])
- plt.show()
- # %% [markdown]
- # ## 2c
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import matplotlib.gridspec as gridspec
- from matplotlib.colors import LinearSegmentedColormap
- # ----------------------------------------------------------------------
- # 1. Load data
- # ----------------------------------------------------------------------
- data = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training_3/Training_Stats/Trained_model_performance_Fig2b.csv')
- # ----------------------------------------------------------------------
- # 2. Clean / normalize model + graph names
- # ----------------------------------------------------------------------
- model_map = {
- 'complex': 'ComplEx',
- 'complexe': 'ComplEx',
- 'simple': 'SimplE',
- 'simpple': 'SimplE',
- 'dismult': 'DisMult',
- 'distmult': 'DisMult',
- 'rescal': 'RESCAL',
- 'rotate': 'RotatE',
- 'transe': 'TransE',
- }
- data['Model'] = data['Model Type'].str.strip().str.lower().map(model_map)
- graph_map = {
- 'EvoAge_121_12M': 'EvoAge',
- 'Aging_121_12M': 'Aging',
- }
- data['Graph'] = data['Graph'].map(graph_map)
- data['metric_type'] = data['metric_type'].str.strip().str.title() # Validation / Testing
- model_order = ['TransE', 'ComplEx', 'RotatE', 'SimplE', 'DisMult', 'RESCAL']
- graph_order = ['Aging', 'EvoAge']
- # ----------------------------------------------------------------------
- # 3. FIGURE 2b — heatmap grid (EvoAge only, MRR + Hit@k)
- # ----------------------------------------------------------------------
- row_order = [('Validation', 'MRR'), ('Validation', 'Hit@1'), ('Validation', 'Hit@3'), ('Validation', 'Hit@10'),
- ('Testing', 'MRR'), ('Testing', 'Hit@1'), ('Testing', 'Hit@3'), ('Testing', 'Hit@10')]
- row_labels = ['MRR', 'hit@1', 'hit@3', 'hit@10', 'MRR', 'hit@1', 'hit@3', 'hit@10']
- n_rows = len(row_order)
- fig = plt.figure(figsize=(10, 4.5))
- gs = gridspec.GridSpec(1, len(model_order) + 1,
- width_ratios=[1] * len(model_order) + [0.12],
- wspace=0.15)
- vmin, vmax = 0.0, 1.0
- # cmap = LinearSegmentedColormap.from_list('original_paper', ['#FCFDDD', '#A9DDD1', '#5EBBA6'])
- # # ----- COLOR GRADIENT: edit these 3 hex codes -----
- # COLOR_LOW = '#FCFDDD' # lowest values (0.0)
- # COLOR_MID = '#A9DDD1' # midpoint (0.5)
- # COLOR_HIGH = '#5EBBA6' # highest values (1.0)
- # COLOR_LOW = '#F7FBFF'
- # COLOR_MID = '#6BAED6'
- # COLOR_HIGH = '#08306B'
- # # cmap = LinearSegmentedColormap.from_list(
- # # 'custom3', [COLOR_LOW, COLOR_MID, COLOR_HIGH]
- # # )
- # cmap = LinearSegmentedColormap.from_list(
- # 'custom3',
- # [(0.0, COLOR_LOW), (0.85, COLOR_MID), (1.0, COLOR_HIGH)] # change 0.5 to shift the middle color
- # )
- # ----- COLOR GRADIENT: 5 stops -----
- C1 = '#F5D49B' # sand
- C2 = '#F0AAAE' # pink
- C3 = '#D98CC8' # mauve
- C4 = '#C063E8' # violet
- C5 = '#9B51E0' # purple
- cmap = LinearSegmentedColormap.from_list('custom5', [C1, C2, C3, C4, C5])
- # # ----- COLOR GRADIENT: 5 stops, edit these hex codes -----
- # C1 = '#FFFFE5' # 0.00 lowest
- # C2 = '#D9F0A3' # 0.25
- # C3 = '#78C679' # 0.50
- # C4 = '#238443' # 0.75
- # C5 = '#004529' # 1.00 highest
- # cmap = LinearSegmentedColormap.from_list('custom5', [C1, C2, C3, C4, C5])
- # ----- COLOR GRADIENT: 5 stops, pastel purple end -----
- C1 = '#F7DFB4' # soft sand
- C2 = '#F2BCBE' # soft pink
- C3 = '#DBA6D2' # soft mauve
- C4 = '#C79BE0' # pastel violet
- C5 = '#AE8CD9' # muted purple
- cmap = LinearSegmentedColormap.from_list('custom5', [
- (0.00, C1),
- (0.20, C2),
- (0.80, C3),
- (0.90, C4),
- (1.00, C5),
- ])
- im = None
- for i, model in enumerate(model_order):
- ax = fig.add_subplot(gs[0, i])
- mat = np.full((n_rows, 1), np.nan)
- for r, (metric_type, col) in enumerate(row_order):
- sub = data[(data['Model'] == model) &
- (data['metric_type'] == metric_type) &
- (data['Graph'] == 'EvoAge')]
- if not sub.empty:
- mat[r, 0] = sub[col].values[0]
- im = ax.imshow(mat, cmap=cmap, vmin=vmin, vmax=vmax, aspect='auto')
- for r in range(n_rows):
- if not np.isnan(mat[r, 0]):
- ax.text(0, r, f'{mat[r, 0]:.2f}', ha='center', va='center', fontsize=8)
- ax.axhline(3.5, color='black', linewidth=1.5)
- ax.set_title(model, fontsize=11, fontweight='bold')
- ax.set_xticks([0])
- ax.set_xticklabels(['EvoAge'], rotation=45, ha='right', fontsize=8)
- ax.set_yticks(range(n_rows))
- ax.set_yticklabels(row_labels if i == 0 else [], fontsize=8)
- ax.tick_params(length=0)
- ax_last = fig.axes[len(model_order) - 1]
- ax_last.text(0.85, 1.5, 'Validation', rotation=-90, va='center', fontsize=9)
- ax_last.text(0.85, 5.5, 'Testing', rotation=-90, va='center', fontsize=9)
- cax = fig.add_subplot(gs[0, -1])
- cbar = fig.colorbar(im, cax=cax)
- cbar.set_label('Score', fontsize=10)
- plt.tight_layout()
- plt.savefig('FIG2/Fig2b_heatmap.svg', dpi=300, bbox_inches='tight')
- plt.show()
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- from matplotlib.colors import LinearSegmentedColormap
- # ----------------------------------------------------------------------
- # 1. Load data
- # ----------------------------------------------------------------------
- EvoAge_6model_Etype = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training_3/Training_Stats/EvoAge_121_12M_64Emb_EdgeType_testing.csv')
- EvoAge_6model_Etype
- # ----------------------------------------------------------------------
- # 2. Order rows by performance (RESCAL, SimplE, DistMult, RotatE, ComplEx, TransE)
- # and select the columns to plot: Hit@1, Hit@3, Hit@10, MRR
- # ----------------------------------------------------------------------
- model_order = ['RESCAL', 'SimplE', 'DistMult', 'RotatE', 'ComplEx', 'TransE']
- df = EvoAge_6model_Etype.set_index('Model Type').loc[model_order]
- cols = ['Hit@1', 'Hit@3', 'Hit@10', 'MRR']
- mat = df[cols].values
- # ----------------------------------------------------------------------
- # 3. FIGURE 2f — edge-type prediction heatmap (Hit@1, Hit@3, Hit@10, MRR)
- # ----------------------------------------------------------------------
- # ----- COLOR GRADIENT: 5 stops, pastel purple end -----
- C1 = '#F7DFB4' # soft sand
- C2 = '#F2BCBE' # soft pink
- C3 = '#DBA6D2' # soft mauve
- C4 = '#C79BE0' # pastel violet
- C5 = '#AE8CD9' # muted purple
- cmap = LinearSegmentedColormap.from_list('custom5', [
- (0.00, C1),
- (0.20, C2),
- (0.80, C3),
- (0.90, C4),
- (1.00, C5),
- ])
- vmin, vmax = 0.0, mat.max()
- fig, ax = plt.subplots(figsize=(5, 4.5))
- im = ax.imshow(mat, cmap=cmap, vmin=vmin, vmax=vmax, aspect='auto')
- # annotate cells
- for r in range(mat.shape[0]):
- for c in range(mat.shape[1]):
- ax.text(c, r, f'{mat[r, c]:.2f}', ha='center', va='center', fontsize=10)
- ax.set_xticks(range(len(cols)))
- ax.set_xticklabels(cols, fontsize=11)
- ax.set_yticks(range(len(model_order)))
- ax.set_yticklabels(model_order, fontsize=11)
- ax.tick_params(length=0)
- # grid lines between cells
- ax.set_xticks(np.arange(-0.5, len(cols), 1), minor=True)
- ax.set_yticks(np.arange(-0.5, len(model_order), 1), minor=True)
- ax.grid(which='minor', color='black', linewidth=1)
- ax.tick_params(which='minor', length=0)
- ax.set_title('Edge-type prediction', fontsize=13)
- cbar = fig.colorbar(im, ax=ax, fraction=0.046, pad=0.04)
- cbar.set_label('score', fontsize=11)
- plt.tight_layout()
- plt.savefig('FIG2/Fig2f_edgetype-prediction.svg', dpi=300, bbox_inches='tight')
- plt.show()
- # %% [markdown]
- # ## 2d
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- plt.rcParams['svg.fonttype'] = 'none'
- # ----------------------------------------------------------------------
- # 1. Load data
- # ----------------------------------------------------------------------
- df = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training_3/Training_Stats/EA_A_B_Rescal_selected_metrics.csv')
- # ----------------------------------------------------------------------
- # 2. Filter: 121_12M ortholog, Testing split only
- # ----------------------------------------------------------------------
- df['Split'] = df['Split'].str.strip().str.title()
- sub = df[(df['Ortholog'] == '121_12M') & (df['Split'] == 'Testing')].copy()
- metrics = ['Hit@1', 'Hit@3', 'Hit@10', 'MRR']
- metric_ticklabels = ['Hit@1', 'Hit@3', 'Hit@10', 'MRR']
- kg_order = ['Aging', 'Biomedical', 'EvoAge']
- kg_colors = {
- 'Aging': '#e17f7f', # salmon/pink
- 'Biomedical': '#7fb3b8', # teal
- 'EvoAge': '#b67fd6', # purple
- }
- # ----------------------------------------------------------------------
- # 3. Build plotting matrix: rows = KG, cols = metrics
- # ----------------------------------------------------------------------
- data = {kg: [sub[sub['KG'] == kg][m].values[0] for m in metrics] for kg in kg_order}
- # ----------------------------------------------------------------------
- # 4. Grouped bar plot
- # ----------------------------------------------------------------------
- x = np.arange(len(metrics))
- n_kg = len(kg_order)
- bar_width = 0.8 / n_kg
- fig, ax = plt.subplots(figsize=(7, 4.5))
- for i, kg in enumerate(kg_order):
- offset = (i - (n_kg - 1) / 2) * bar_width
- bars = ax.bar(x + offset, data[kg], width=bar_width,
- color=kg_colors[kg], edgecolor='black', linewidth=0.6,
- label=kg)
- for rect, val in zip(bars, data[kg]):
- ax.text(rect.get_x() + rect.get_width() / 2, rect.get_height() + 0.01,
- f'{val:.2f}', ha='center', va='bottom', fontsize=8)
- ax.set_xticks(x)
- ax.set_xticklabels(metric_ticklabels, fontsize=10)
- ax.set_ylabel('Score', fontsize=10)
- ax.set_ylim(0, 1.05)
- ax.set_title('RESCAL entity (Testing)', fontsize=11)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- ax.legend(frameon=False, fontsize=9, loc='upper right', bbox_to_anchor=(1.15, 1.0))
- plt.tight_layout()
- plt.savefig('FIG2/Fig2_121_12M_testing_barplot_rescal_entity.svg', dpi=300, bbox_inches='tight')
- plt.show()
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import matplotlib as mpl
- from matplotlib.patches import Rectangle
- from matplotlib.lines import Line2D
- mpl.rcParams['svg.fonttype'] = 'none'
- # ----------------------------------------------------------------------
- # 1. Load data
- # ----------------------------------------------------------------------
- EvoAge_Aging_biomed_edgeType = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training_3/Training_Stats/RESCAL_all_KGtype_edgetype_results.csv')
- df = EvoAge_Aging_biomed_edgeType.copy()
- def get_category(name):
- if name.startswith('Aging'):
- return 'Aging'
- elif name.startswith('Biomedical'):
- return 'Biomedical'
- elif name.startswith('EvoAge'):
- return 'EvoAge'
- def get_ortholog_type(name):
- return '1:1' if '1to1' in name else '1:N'
- df['Category'] = df['Dataset'].apply(get_category)
- df['OrthologType'] = df['Dataset'].apply(get_ortholog_type)
- cols = ['Hit@1', 'Hit@3', 'Hit@10', 'MRR']
- row_order = ['1:N']
- # ----------------------------------------------------------------------
- # 2. Flat category colours -- sampled straight from your legend snapshot
- # ----------------------------------------------------------------------
- panel_colors = {
- "Aging": "#e69797",
- "Biomedical": "#80b4b9",
- "EvoAge": "#b780d6",
- }
- panel_titles = {'Aging': 'Aging', 'Biomedical': 'Biomedical', 'EvoAge': 'EvoAge'}
- panel_letters = {'Aging': 'b', 'Biomedical': 'c', 'EvoAge': 'd'}
- # 1:1 = solid full colour, 1:N = same hue but lighter (alpha)
- bar_alpha = {'1:1': 1.0, '1:N': 0.5}
- row_swatch_colors = ['#8b96c9', '#e0a377'] # legend swatches (kept from original style)
- # ----------------------------------------------------------------------
- # 3. Build one combined figure with 3 grouped-bar panels
- # ----------------------------------------------------------------------
- fig, axes = plt.subplots(1, 3, figsize=(13, 3.8))
- plt.subplots_adjust(wspace=0.45, top=0.72, bottom=0.28)
- x = np.arange(len(cols))
- bar_width = 0.5
- for ax, category in zip(axes, ['Aging', 'Biomedical', 'EvoAge']):
- sub = df[df['Category'] == category].set_index('OrthologType').loc[row_order]
- color = panel_colors[category]
- vals = sub.loc['1:N', cols].values.astype(float)
- ax.bar(x, vals, width=bar_width,
- color=color, alpha=bar_alpha['1:N'],
- edgecolor='black', linewidth=0.8)
- for xi, v in zip(x, vals):
- ax.text(xi, v + 0.015, f'{v:.2f}', ha='center', va='bottom', fontsize=9)
- ax.set_xticks(x)
- ax.set_xticklabels(['Hit@1', 'Hit@3', 'Hit@10', 'MRR'], fontsize=10)
- ax.set_ylim(0, 1.05)
- ax.set_yticks([0.0, 0.2, 0.4, 0.6, 0.8, 1.0])
- ax.tick_params(axis='y', labelsize=9)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- ax.set_title(panel_titles[category], fontsize=13, pad=10)
- # panel letter, top-left of panel
- ax.text(-0.12, 1.08, panel_letters[category], fontsize=16, fontweight='bold',
- ha='left', va='bottom', transform=ax.transAxes, clip_on=False)
- # top overall title with flanking lines
- fig.text(0.5, 0.92, 'RESCAL edgetype prediction', ha='center', va='center',
- fontsize=14, fontweight='bold')
- fig.add_artist(Line2D([0.06, 0.42], [0.93, 0.93], color='black', lw=1.3, transform=fig.transFigure))
- fig.add_artist(Line2D([0.58, 0.94], [0.93, 0.93], color='black', lw=1.3, transform=fig.transFigure))
- # bottom label: One-to-One plus One-to-Many (single group now, no legend split needed)
- legend_y = 0.06
- fig.text(0.5, legend_y, 'One-to-One plus One-to-Many', ha='center', va='center', fontsize=11)
- fig.add_artist(Line2D([0.06, 0.40], [legend_y, legend_y], color='black', lw=1.3, transform=fig.transFigure))
- fig.add_artist(Line2D([0.60, 0.94], [legend_y, legend_y], color='black', lw=1.3, transform=fig.transFigure))
- plt.savefig('FIG2/Fig2_bcd_RESCAL_bars.svg', bbox_inches='tight')
- plt.savefig('FIG2/Fig2_bcd_RESCAL_bars.png', dpi=300, bbox_inches='tight')
- plt.show()
- print("done")
- # %% [markdown]
- # ## 2e
- # %%
- import pandas as pd
- import matplotlib.pyplot as plt
- import numpy as np
- # --------------------------------------------------
- # Load data
- # --------------------------------------------------
- df = pd.read_csv("/storage/Arushi/090526_EvoAge/kg_formation/training_3/InAccuracy_format/combined_classification_accuracy_R4.csv")
- graphs = [
- ("Aging_121_12M", "Aging", "#e69797"),
- ("Biomedical_121_12M", "Biomedical", "#80b4b9"),
- ("EvoAge_121_12M", "EvoAge", "#b780d6"),
- ]
- metrics = ["Accuracy", "Precision", "Recall", "F1_Score", "Mean_AUC"]
- labels = ["Accuracy", "Precision", "Recall", "F1", "AUC"]
- plt.rcParams.update({
- "font.family": "DejaVu Sans",
- "font.size": 12,
- "axes.linewidth": 1.2,
- "svg.fonttype": "none" # keeps text editable in Illustrator/Inkscape
- })
- # --------------------------------------------------
- # Create one figure with 3 panels
- # --------------------------------------------------
- fig, axes = plt.subplots(1, 3, figsize=(15, 4.5), sharey=True)
- for ax, (graph, title, color) in zip(axes, graphs):
- row = df[df["Graph"] == graph].iloc[0]
- values = row[metrics].astype(float).values
- bars = ax.bar(
- np.arange(len(metrics)),
- values,
- width=0.62,
- color=color,
- edgecolor="black",
- linewidth=0.8
- )
- # value labels
- for bar, val in zip(bars, values):
- ax.text(
- bar.get_x() + bar.get_width()/2,
- val + 0.008,
- f"{val:.2f}",
- ha="center",
- va="bottom",
- fontsize=9
- )
- ax.set_xticks(np.arange(len(metrics)))
- ax.set_xticklabels(labels, fontsize=10)
- ax.set_title(title, fontsize=15, weight="bold")
- ax.grid(axis="y", linestyle="--", alpha=0.3)
- ax.set_axisbelow(True)
- ax.spines["top"].set_visible(False)
- ax.spines["right"].set_visible(False)
- axes[0].set_ylabel("Score", fontsize=12)
- axes[0].set_ylim(0.60, 1.05)
- # Optional overall title
- # fig.suptitle("Node Classification Performance (1:1 + 1:N)", fontsize=16)
- plt.tight_layout()
- # --------------------------------------------------
- # Save ONE SVG and ONE PNG
- # --------------------------------------------------
- plt.savefig(
- "FIG2/triple_classification_KG_Accuracy_1to1plus1toMany.svg",
- format="svg",
- bbox_inches="tight"
- )
- plt.savefig(
- "FIG2/_TRIPLE_CLASSIFICATION_KG_Accuracy_1to1plus1toMany.png",
- dpi=600,
- bbox_inches="tight"
- )
- plt.show()
- # %% [markdown]
- # ## 2f
- # %%
- # Edge_type = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training__2/Shuffled_EvoAge_testing/Store_House/shuffled_test_sets/rescal_shuffled_metrics.csv')
- # Edge_type
- #!/usr/bin/env python3
- """
- Monte Carlo P-Value Analysis — Radial (Spider) Plot
- Compare real EvoAge test set performance vs 30 shuffled test sets.
- Produces a single radial/spider chart (matching the reference figure):
- - Red polygon = real EvoAge test performance
- - Blue polygon = mean of 30 shuffled null-hypothesis baselines
- - Value labels on every spoke
- - "Statistical Significance" box (per-metric Monte Carlo p-value + marker)
- - "Interpretation" box
- Output is saved as a PDF (vector, publication-ready).
- """
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- from scipy import stats
- plt.rcParams.update({
- "font.family": "DejaVu Sans",
- "svg.fonttype": "none"
- })
- # ---- Your real test set results ----
- REAL_RESULTS = {
- 'HITS@1': 0.810538,
- 'HITS@3': 0.924235,
- 'HITS@10': 0.997215,
- 'MR': 1.520212,
- 'MRR': 0.875109
- }
- # ---- Load your shuffled results ----
- results_df = pd.read_csv(
- '/storage/Arushi/090526_EvoAge/kg_formation/training_3/Shuffled_EvoAge_testing/Store_House/shuffled_test_sets/rescal_shuffled_metrics.csv'
- )
- print("=" * 70)
- print("MONTE CARLO P-VALUE ANALYSIS")
- print("Real Test Set vs 30 Shuffled Test Sets")
- print("=" * 70)
- print(f"\n[1/3] Loaded {len(results_df)} shuffled test set results")
- print(results_df[['shuffle_id', 'MRR', 'HITS@1', 'HITS@10']].head(10))
- # Order controls the layout of the radial chart (clockwise from the top)
- metrics = ['MRR', 'HITS@1', 'HITS@3', 'HITS@10']
- # For MRR, HITS@K: higher = better -> test if real is significantly HIGHER
- # For MR: lower = better -> test if real is significantly LOWER
- higher_is_better = {'MRR': True, 'HITS@1': True, 'HITS@3': True, 'HITS@10': True}
- # ---- Compute p-values ----
- print("\n[2/3] Computing Monte Carlo p-values...")
- summary_rows = []
- for metric in metrics:
- shuffled_vals = results_df[metric].dropna().values
- real_val = REAL_RESULTS[metric]
- n = len(shuffled_vals)
- shuffled_mean = shuffled_vals.mean()
- shuffled_std = shuffled_vals.std(ddof=1)
- shuffled_min = shuffled_vals.min()
- shuffled_max = shuffled_vals.max()
- # ---- MONTE CARLO P-VALUE (empirical) ----
- if higher_is_better[metric]:
- n_extreme = np.sum(shuffled_vals >= real_val)
- else:
- n_extreme = np.sum(shuffled_vals <= real_val)
- # Empirical p-value with +1/+1 correction (standard in permutation testing)
- monte_carlo_p = (n_extreme + 1) / (n + 1)
- # ---- Z-TEST P-VALUE (parametric, for reference) ----
- z_score = (real_val - shuffled_mean) / shuffled_std if shuffled_std > 0 else np.nan
- if higher_is_better[metric]:
- z_p = stats.norm.sf(z_score)
- else:
- z_p = stats.norm.cdf(z_score)
- # Cohen's d (effect size)
- cohens_d = (real_val - shuffled_mean) / shuffled_std if shuffled_std > 0 else np.nan
- summary_rows.append({
- 'metric': metric,
- 'real_value': real_val,
- 'shuffled_mean': shuffled_mean,
- 'shuffled_std': shuffled_std,
- 'shuffled_min': shuffled_min,
- 'shuffled_max': shuffled_max,
- 'n_shuffles': n,
- 'n_extreme': n_extreme,
- 'z_score': z_score,
- 'cohens_d': cohens_d,
- 'p_monte_carlo': monte_carlo_p,
- 'p_ztest': z_p
- })
- summary_df = pd.DataFrame(summary_rows)
- pd.set_option('display.float_format', lambda x: f'{x:.6g}')
- pd.set_option('display.max_columns', None)
- pd.set_option('display.width', None)
- print("\n" + summary_df.to_string(index=False))
- # ---- Save results table ----
- output_dir = 'FIG2/'
- summary_df.to_csv(f'{output_dir}/monte_carlo_pvalues_vs_shuffled.csv', index=False)
- print(f"\n✓ Saved to: {output_dir}/monte_carlo_pvalues_vs_shuffled.csv")
- # ---- Radial (spider) plot ----
- print("\n[3/3] Creating radial visualization...")
- def sig_marker(p):
- if p < 0.001:
- return '***'
- if p < 0.01:
- return '**'
- if p < 0.05:
- return '*'
- return 'n.s.'
- real_vals = [summary_df.loc[summary_df['metric'] == m, 'real_value'].values[0] for m in metrics]
- shuf_vals = [summary_df.loc[summary_df['metric'] == m, 'shuffled_mean'].values[0] for m in metrics]
- N = len(metrics)
- angles = np.linspace(0, 2 * np.pi, N, endpoint=False).tolist()
- real_plot = real_vals + real_vals[:1]
- shuf_plot = shuf_vals + shuf_vals[:1]
- angles_plot = angles + angles[:1]
- fig = plt.figure(figsize=(11, 10))
- ax = fig.add_subplot(111, polar=True)
- ax.set_theta_zero_location('N') # first metric (MRR) at the top
- ax.set_theta_direction(-1) # go clockwise
- ax.set_xticks(angles)
- ax.set_xticklabels(metrics, fontsize=12, fontweight='bold')
- max_r = max(real_vals + shuf_vals) * 1.25
- ax.set_ylim(0, max_r)
- ax.set_rgrids(np.linspace(max_r / 5, max_r, 5), angle=0, fontsize=8)
- ax.grid(alpha=0.4, linestyle='--')
- RED = '#d62728'
- BLUE = '#4a78c9'
- # Real EvoAge performance
- ax.plot(angles_plot, real_plot, color=RED, linewidth=2.2, label='EvoAge test')
- ax.fill(angles_plot, real_plot, color=RED, alpha=0.15)
- ax.scatter(angles, real_vals, color=RED, s=45, zorder=5)
- # Shuffled null baseline (mean of 30 shuffles)
- ax.plot(angles_plot, shuf_plot, color=BLUE, linewidth=2.2, label='EvoAge Shuffled test')
- ax.fill(angles_plot, shuf_plot, color=BLUE, alpha=0.12)
- ax.scatter(angles, shuf_vals, color=BLUE, s=45, zorder=5)
- # Value labels
- for ang, val in zip(angles, real_vals):
- ax.annotate(f'{val:.4f}', xy=(ang, val), xytext=(ang, val + max_r * 0.03),
- ha='center', va='bottom', fontsize=9, fontweight='bold', color=RED,
- bbox=dict(boxstyle='round,pad=0.25', fc='white', ec=RED, lw=1))
- for ang, val in zip(angles, shuf_vals):
- ax.annotate(f'{val:.4f}', xy=(ang, val), xytext=(ang, val - max_r * 0.03),
- ha='center', va='top', fontsize=8, style='italic', color=BLUE,
- bbox=dict(boxstyle='round,pad=0.2', fc='white', ec=BLUE, lw=0.8))
- ax.set_title(
- 'RESCAL KG Embedding: Monte Carlo Null Hypothesis Testing\n'
- 'Real Performance vs 30 Shuffled Test Set Baselines (Radial View)',
- fontsize=13, fontweight='bold', pad=30
- )
- ax.legend(loc='lower right', bbox_to_anchor=(1.3, -0.05), fontsize=12, frameon=False)
- # Statistical significance box
- sig_lines = "\n".join(
- f"{row['metric']}: {sig_marker(row['p_monte_carlo'])} (p={row['p_monte_carlo']:.3g})"
- for _, row in summary_df.iterrows()
- )
- fig.text(0.03, 0.35,
- f"Statistical Significance:\n\n{sig_lines}",
- fontsize=9, family='monospace',
- bbox=dict(boxstyle='round,pad=0.6', fc='#fdf6d8', ec='#c9b95a', lw=1))
- # Interpretation box
- interp_text = (
- "Interpretation:\n"
- "- Red polygon = Real EvoAge KG performance\n"
- "- Blue polygon = Shuffled null baseline (mean)\n"
- "- Larger red area beyond blue = significant gain\n"
- "- Values labeled on each spoke"
- )
- fig.text(0.03, 0.12, interp_text, fontsize=9,
- bbox=dict(boxstyle='round,pad=0.6', fc='#e4f2fb', ec='#7fb3d9', lw=1))
- fig.tight_layout()
- pdf_path = f'{output_dir}/monte_carlo_radial_comparison.pdf'
- fig.savefig(pdf_path, format='pdf', bbox_inches='tight')
- print(f"✓ Saved radial plot to: {pdf_path}")
- plt.show()
- plt.close(fig)
- # ---- Final interpretation (console) ----
- print("\n" + "=" * 70)
- print("INTERPRETATION")
- print("=" * 70)
- for _, row in summary_df.iterrows():
- metric = row['metric']
- real = row['real_value']
- shuffled_mean = row['shuffled_mean']
- p_val = row['p_monte_carlo']
- cohens_d = row['cohens_d']
- if p_val < 0.001:
- sig_text = "*** HIGHLY SIGNIFICANT"
- elif p_val < 0.01:
- sig_text = "** VERY SIGNIFICANT"
- elif p_val < 0.05:
- sig_text = "* SIGNIFICANT"
- else:
- sig_text = "n.s. NOT SIGNIFICANT"
- ratio = real / shuffled_mean if shuffled_mean != 0 else np.inf
- print(f"\n{metric:10s}")
- print(f" Real: {real:.6f}")
- print(f" Shuffled mean: {shuffled_mean:.6f} (ratio: {ratio:.2f}x)")
- print(f" p-value: {p_val:.6f} {sig_text}")
- print(f" Effect size: Cohen's d = {cohens_d:+.3f}")
- print("\n" + "=" * 70)
- # %% [markdown]
- # ## 2h
- # %%
- import pandas as pd
- import matplotlib.pyplot as plt
- import numpy as np
- plt.rcParams.update({
- "font.family": "DejaVu Sans",
- "svg.fonttype": "none"
- })
- k = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/building_evoage_with_1_percent_species_testsplit/1_per_test_triple_count.csv')
- print(k.columns.tolist())
- # Color scheme matching original figure
- colors = {
- 'Human': '#A98BC9',
- 'Mouse': '#8AD3E4',
- 'Celegans': '#C9C97E',
- 'Drosophila': '#8FCB7E',
- 'Zebrafish': '#E89BC0',
- 'Yeast': '#C9A98B',
- 'CrossSpecies': '#F7DADF',
- }
- order = ['Human', 'Mouse', 'Celegans', 'Drosophila', 'Zebrafish', 'Yeast', 'CrossSpecies']
- # ---------- Panel K: donut chart of heldout data (log-scaled wedges, real labels) ----------
- k = k.set_index('Species').loc[order].reset_index()
- # Log-transform wedge sizes so small species are still visible,
- # but keep the real counts as text labels.
- log_vals = np.log10(k['Line_Count'])
- log_vals = log_vals - log_vals.min() + 1 # shift so smallest still gets a visible wedge
- fig, ax = plt.subplots(figsize=(6, 6))
- wedges, _ = ax.pie(
- log_vals,
- colors=[colors[s] for s in k['Species']],
- startangle=90,
- counterclock=False,
- wedgeprops=dict(width=0.35, edgecolor='white', linewidth=1.5)
- )
- total = k['Line_Count'].sum()
- ax.text(0, 0.05, 'total', ha='center', va='center', fontsize=13, fontweight='bold')
- ax.text(0, -0.08, f'{total:,}', ha='center', va='center', fontsize=12)
- # Labels read back from each wedge's true midpoint angle -> can never drift onto the wrong wedge
- for w, (_, row) in zip(wedges, k.iterrows()):
- ang = np.radians((w.theta1 + w.theta2) / 2)
- x = 1.3 * np.cos(ang)
- y = 1.3 * np.sin(ang)
- ax.text(x, y, f"{row['Species']}\n{row['Line_Count']:,}",
- ha='center', va='center', fontsize=8)
- ax.set_title('heldout data (wedges log-scaled; labels = actual counts)',
- fontsize=12, fontweight='bold', loc='left')
- plt.tight_layout()
- plt.savefig('FIG2/fig2n.pdf')
- plt.show()
- plt.close()
- k.head()
- # %% [markdown]
- # ## 2i
- # %%
- import pandas as pd
- import matplotlib.pyplot as plt
- import numpy as np
- plt.rcParams.update({
- "font.family": "DejaVu Sans",
- "svg.fonttype": "none"
- })
- # ================================================================
- # CANONICAL PALETTE (matches icon legend)
- # ================================================================
- colors = {
- 'Human': '#A98BC9',
- 'Mouse': '#8AD3E4',
- 'Celegans': '#C9C97E',
- 'Drosophila': '#8FCB7E',
- 'Zebrafish': '#E89BC0',
- 'Yeast': '#C9A98B',
- 'CrossSpecies': '#F7DADF',
- }
- order = ['Human', 'Mouse', 'Celegans', 'Drosophila', 'Zebrafish', 'Yeast', 'CrossSpecies']
- # ================================================================
- # FIG 2N — heldout-data donut
- # ================================================================
- def make_donut():
- k = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/'
- 'building_evoage_with_1_percent_species_testsplit/1_per_test_triple_count.csv')
- k = k.set_index('Species').loc[order].reset_index()
- log_vals = np.log10(k['Line_Count'])
- log_vals = log_vals - log_vals.min() + 1 # smallest wedge stays visible
- fig, ax = plt.subplots(figsize=(6, 6))
- wedges, _ = ax.pie(
- log_vals,
- colors=[colors[s] for s in k['Species']],
- startangle=90,
- counterclock=False,
- wedgeprops=dict(width=0.35, edgecolor='white', linewidth=1.5)
- )
- total = k['Line_Count'].sum()
- ax.text(0, 0.05, 'total', ha='center', va='center', fontsize=13, fontweight='bold')
- ax.text(0, -0.08, f'{total:,}', ha='center', va='center', fontsize=12)
- # labels read from each wedge's true midpoint -> never drift onto wrong wedge
- for w, (_, row) in zip(wedges, k.iterrows()):
- ang = np.radians((w.theta1 + w.theta2) / 2)
- x, y = 1.3 * np.cos(ang), 1.3 * np.sin(ang)
- ax.text(x, y, f"{row['Species']}\n{row['Line_Count']:,}",
- ha='center', va='center', fontsize=8)
- ax.set_title('heldout data (wedges log-scaled; labels = actual counts)',
- fontsize=12, fontweight='bold', loc='left')
- plt.tight_layout()
- plt.savefig('FIG2/fig2n.pdf', bbox_inches='tight')
- plt.close()
- # ================================================================
- # FIG 2O — link-prediction hit@k grouped bars
- # ================================================================
- def make_linkpred():
- l = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training_3/'
- '1_per_species_test_EvoAge_121_12M/rescal_1_percent_test_linkprediction_results.csv')
- l['TestSet'] = l['TestSet'].replace({'Unknown_0': 'Human', 'Unknown_6': 'CrossSpecies'})
- l = l.set_index('TestSet').loc[order].reset_index() # ordered; index i == order[i]
- hits = ['Hits@1', 'Hits@3', 'Hits@10']
- x_labels = ['1', '3', '10']
- n_species = len(order)
- width = 0.8 / n_species
- x = np.arange(len(hits))
- fig, ax = plt.subplots(figsize=(6, 5))
- for i, sp in enumerate(order):
- vals = l.iloc[i][hits].values.astype(float) # by position -> no filter mismatch
- ax.bar(x + (i - n_species / 2) * width + width / 2, vals,
- width=width, color=colors[sp], label=sp)
- ax.set_ylim(0, 1.05)
- ax.set_xticks(x)
- ax.set_xticklabels(x_labels)
- ax.set_xlabel('hit@')
- ax.set_ylabel('score')
- ax.set_title('prediction (link)', fontsize=14, fontweight='bold')
- ax.legend(bbox_to_anchor=(1.01, 1), loc='upper left', fontsize=8, frameon=False)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- plt.tight_layout()
- plt.savefig('FIG2/fig2o_new.pdf', bbox_inches='tight')
- plt.show()
- plt.close()
- if __name__ == '__main__':
- make_donut()
- make_linkpred()
- print('done')
- # %%
- import pandas as pd
- import matplotlib.pyplot as plt
- import numpy as np
- plt.rcParams.update({
- "font.family": "DejaVu Sans",
- "svg.fonttype": "none"
- })
- # ================================================================
- # CANONICAL PALETTE (matches icon legend)
- # ================================================================
- colors = {
- 'Human': '#A98BC9',
- 'Mouse': '#8AD3E4',
- 'Celegans': '#C9C97E',
- 'Drosophila': '#8FCB7E',
- 'Zebrafish': '#E89BC0',
- 'Yeast': '#C9A98B',
- 'CrossSpecies': '#F7DADF',
- }
- order = ['Human', 'Mouse', 'Celegans', 'Drosophila', 'Zebrafish', 'Yeast', 'CrossSpecies']
- # ================================================================
- # FIG 2N — heldout-data donut
- # ================================================================
- def make_donut():
- k = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/'
- 'building_evoage_with_1_percent_species_testsplit/1_per_test_triple_count.csv')
- k = k.set_index('Species').loc[order].reset_index()
- log_vals = np.log10(k['Line_Count'])
- log_vals = log_vals - log_vals.min() + 1 # smallest wedge stays visible
- fig, ax = plt.subplots(figsize=(6, 6))
- wedges, _ = ax.pie(
- log_vals,
- colors=[colors[s] for s in k['Species']],
- startangle=90,
- counterclock=False,
- wedgeprops=dict(width=0.35, edgecolor='white', linewidth=1.5)
- )
- total = k['Line_Count'].sum()
- ax.text(0, 0.05, 'total', ha='center', va='center', fontsize=13, fontweight='bold')
- ax.text(0, -0.08, f'{total:,}', ha='center', va='center', fontsize=12)
- # labels read from each wedge's true midpoint -> never drift onto wrong wedge
- for w, (_, row) in zip(wedges, k.iterrows()):
- ang = np.radians((w.theta1 + w.theta2) / 2)
- x, y = 1.3 * np.cos(ang), 1.3 * np.sin(ang)
- ax.text(x, y, f"{row['Species']}\n{row['Line_Count']:,}",
- ha='center', va='center', fontsize=8)
- ax.set_title('heldout data (wedges log-scaled; labels = actual counts)',
- fontsize=12, fontweight='bold', loc='left')
- plt.tight_layout()
- plt.savefig('FIG2/fig2n.pdf', bbox_inches='tight')
- plt.close()
- # ================================================================
- # FIG 2O — link-prediction hit@k grouped bars
- # ================================================================
- def make_linkpred():
- l = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training_3/'
- '1_per_species_test_EvoAge_121_12M/rescal_1_percent_test_linkprediction_results.csv')
- l['TestSet'] = l['TestSet'].replace({'Unknown_0': 'Human', 'Unknown_6': 'CrossSpecies'})
- l = l.set_index('TestSet').loc[order].reset_index() # ordered; index i == order[i]
- hits = ['Hits@1', 'Hits@3', 'Hits@10']
- x_labels = ['1', '3', '10']
- n_species = len(order)
- width = 0.8 / n_species
- x = np.arange(len(hits))
- fig, ax = plt.subplots(figsize=(6, 5))
- for i, sp in enumerate(order):
- vals = l.iloc[i][hits].values.astype(float) # by position -> no filter mismatch
- ax.bar(x + (i - n_species / 2) * width + width / 2, vals,
- width=width, color=colors[sp], label=sp)
- ax.set_ylim(0, 1.05)
- ax.set_xticks(x)
- ax.set_xticklabels(x_labels)
- ax.set_xlabel('hit@')
- ax.set_ylabel('score')
- ax.set_title('prediction (link)', fontsize=14, fontweight='bold')
- ax.legend(bbox_to_anchor=(1.01, 1), loc='upper left', fontsize=8, frameon=False)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- plt.tight_layout()
- plt.savefig('FIG2/fig2o_new.pdf', bbox_inches='tight')
- plt.close()
- # ================================================================
- # FIG 2P — edge-type prediction hit@k line plot
- # ================================================================
- def make_edgetype():
- m = pd.read_csv('/storage/Arushi/090526_EvoAge/kg_formation/training_3/'
- '1_per_species_test_EvoAge_121_12M/rescal_1_percent_test_edgeType_results.csv')
- m['TestSet'] = m['TestSet'].replace({'Unknown_0': 'Human', 'Unknown_6': 'CrossSpecies'})
- print(m['TestSet'].unique())
- hits = ['Hits@1', 'Hits@3', 'Hits@10']
- available = m['TestSet'].unique()
- order_m = [s for s in order if s in available] # keep canonical order, drop missing
- fig, ax = plt.subplots(figsize=(6, 5))
- x_pos = [0, 1, 2] # evenly spaced positions so '3' sits centered, not skewed toward '1'
- for sp in order_m:
- vals = m.loc[m['TestSet'] == sp, hits].values.flatten().astype(float)
- ax.plot(x_pos, vals, marker='o', color=colors[sp], label=sp, linewidth=2)
- ax.set_ylim(0, 1.05)
- ax.set_xticks(x_pos)
- ax.set_xticklabels(['1', '3', '10'])
- ax.set_xlabel('hit@')
- ax.set_ylabel('score')
- ax.set_title('prediction (edge type)', fontsize=14, fontweight='bold')
- ax.legend(bbox_to_anchor=(1.01, 1), loc='upper left', fontsize=8, frameon=False)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- plt.tight_layout()
- plt.savefig('FIG2/fig2p_new.pdf', bbox_inches='tight')
- plt.show()
- plt.close()
- if __name__ == '__main__':
- make_donut()
- make_linkpred()
- make_edgetype()
- print('done')
- # %% [markdown]
- # ## 2j
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- plt.rcParams.update({
- "font.family": "DejaVu Sans",
- "svg.fonttype": "none"
- })
- # ----------------------------------------------------------------------
- # 1. Load data
- # ----------------------------------------------------------------------
- dataframe = pd.read_csv(
- "/storage/Arushi/090526_EvoAge/kg_formation/training_3/Agingonly_testing_data/Aging_Only_Testingdata_Evaluated_onall_rescal_trainedKG's_summary.csv"
- )
- # ----------------------------------------------------------------------
- # 2. Keep only 1:N (121_12M)
- # ----------------------------------------------------------------------
- dataframe = dataframe[dataframe["Ortholog"] == "121_12M"].copy()
- ortholog_label = {"121_12M": "1:N"}
- dataframe["OrthologLabel"] = dataframe["Ortholog"].map(ortholog_label)
- dataframe["Condition"] = dataframe["Graph"] + " (" + dataframe["OrthologLabel"] + ")"
- # Order
- condition_order = [
- "Aging -> Aging (1:N)",
- "Aging -> EvoAge (1:N)",
- ]
- dataframe["Condition"] = pd.Categorical(
- dataframe["Condition"],
- categories=condition_order,
- ordered=True,
- )
- dataframe = dataframe.sort_values("Condition")
- # Colours
- condition_colors = {
- "Aging -> Aging (1:N)": "#e07a7a",
- "Aging -> EvoAge (1:N)": "#4f8fd9",
- }
- metrics = ["Hit@1", "Hit@3", "Hit@10", "MRR"]
- # ----------------------------------------------------------------------
- # 3. Grouped bar chart
- # ----------------------------------------------------------------------
- fig, ax = plt.subplots(figsize=(6, 5))
- n_conditions = len(condition_order)
- width = 0.35
- x = np.arange(len(metrics))
- for i, cond in enumerate(condition_order):
- row = dataframe[dataframe["Condition"] == cond]
- if row.empty:
- continue
- vals = row[metrics].values.flatten()
- offset = (i - (n_conditions - 1) / 2) * width
- ax.bar(
- x + offset,
- vals,
- width,
- label=cond,
- color=condition_colors[cond],
- )
- ax.set_xticks(x)
- ax.set_xticklabels(metrics, fontsize=11)
- ax.set_ylabel("Score", fontsize=12)
- ax.set_ylim(0, 1.05)
- ax.legend(frameon=False, fontsize=10)
- ax.spines["top"].set_visible(False)
- ax.spines["right"].set_visible(False)
- plt.tight_layout()
- plt.savefig("FIG2/2oAging_crossKG_barplot_1N_only.pdf", dpi=300, bbox_inches="tight")
- plt.show()
- # %% [markdown]
- # ## Main Figure 4
- # %% [markdown]
- # ## 4b
- # %%
- from __future__ import annotations
- %matplotlib inline
- import csv
- import re
- import shutil
- import urllib.parse
- import zipfile
- from collections import Counter, defaultdict
- from pathlib import Path
- import numpy as np
- import matplotlib.pyplot as plt
- from matplotlib.colors import LinearSegmentedColormap
- from matplotlib.backends.backend_pdf import PdfPages
- # ============================================================
- # Configuration
- # ============================================================
- INPUT = Path('/storage/Arushi/090526_EvoAge/multiagent_hypo/evoage_100_right_inverse/all_LLMs_hypothesis_response.csv')
- ROOT = Path('/storage/Arushi/090526_EvoAge/multiagent_hypo/evoage_100_right_inverse/My_final_analysis/My_analysis_doi_paired_retention_plots')
- MASTER_PDF = ROOT / 'retention_pie_donut_plots.pdf'
- ALL_MODELS = ['EvoAge', 'medgemma', 'BioMistral', 'BioMistralFinetuned']
- DISPLAY = {
- 'EvoAge': 'EvoAge',
- 'medgemma': 'MedGemma',
- 'BioMistral': 'BioMistral',
- 'BioMistralFinetuned': 'BioMistral fine-tuned',
- }
- PRIMARY_MODELS = ['EvoAge', 'medgemma', 'BioMistral']
- ALT_MODELS = ['EvoAge', 'BioMistral', 'BioMistralFinetuned']
- TYPES = ['Right', 'Inverse']
- VERDICTS = ['no_support', 'weak_support', 'partial_support', 'support', 'strong_support']
- VERDICT_LABELS = ['No support', 'Weak support', 'Partial support', 'Support', 'Strong support']
- SCORE = {'no_support': 0, 'weak_support': 1, 'partial_support': 2, 'support': 3, 'strong_support': 4}
- UNKNOWN = {'unknown', 'missing', 'unknown/missing', 'unknown / missing', ''}
- MODEL_COLORS = {'EvoAge': '#4C72B0', 'medgemma': '#DD8452', 'BioMistral': '#55A868',
- 'BioMistralFinetuned': '#937860'}
- KEPT_COLOR = '#4C9F70'
- REMOVED_COLOR = '#C44E52'
- MULTI_COLOR = '#8172B3'
- BAR_COLOR = '#1f77b4'
- GRADIENT_ENDPOINTS = ['#1a9850', '#d61f96']
- _ordinal_cmap = LinearSegmentedColormap.from_list('green_to_magenta', GRADIENT_ENDPOINTS)
- VERDICT_COLORS = [_ordinal_cmap(i / (len(VERDICTS) - 1)) for i in range(len(VERDICTS))]
- plt.rcParams.update({
- 'font.family': 'DejaVu Sans',
- 'font.size': 9,
- 'axes.titlesize': 11,
- 'axes.labelsize': 9,
- 'legend.fontsize': 8,
- 'axes.linewidth': 0.8,
- 'patch.linewidth': 0.7,
- 'axes.spines.top': False,
- 'axes.spines.right': False,
- })
- # ============================================================
- # Helper Functions
- # ============================================================
- def normalize_doi(value: str) -> str:
- value = (value or '').strip().lower()
- value = re.sub(r'^https?://(dx\.)?doi\.org/', '', value)
- value = re.sub(r'^doi:\s*', '', value)
- value = urllib.parse.unquote(value).strip().rstrip(' .;,')
- return value
- def normalize_verdict(value: str) -> str:
- value = (value or '').strip().lower().replace('-', '_').replace(' ', '_')
- value = re.sub(r'_+', '_', value)
- aliases = {'nosupport': 'no_support', 'partialsupport': 'partial_support',
- 'weaksupport': 'weak_support', 'strongsupport': 'strong_support',
- 'unknown_missing': 'unknown'}
- return aliases.get(value, value)
- def footer(fig, text):
- fig.text(0.01, 0.01, text, ha='left', va='bottom', fontsize=6.5)
- def pct_count_autopct(total: int):
- def _fmt(pct):
- count = int(round(pct * total / 100.0))
- return f'{pct:.1f}%\n(n={count})'
- return _fmt
- # ============================================================
- # Load and Process Data
- # ============================================================
- with INPUT.open(newline='', encoding='utf-8-sig') as handle:
- reader = csv.DictReader(handle)
- raw = []
- for row in reader:
- model = row['model'].strip()
- if model not in ALL_MODELS:
- continue
- raw.append({
- 'doi': normalize_doi(row['DOI']),
- 'model': model,
- 'hypothesis_type': row['hypothesis_type'].strip().title(),
- 'verdict': normalize_verdict(row['verdict']),
- })
- by_doi = defaultdict(list)
- for row in raw:
- by_doi[row['doi']].append(row)
- dois = sorted(by_doi)
- lookup_all = {(r['doi'], r['model'], r['hypothesis_type']): r for r in raw}
- def scores(model: str, direction: str, doi_list) -> np.ndarray:
- return np.array([SCORE[lookup_all[(doi, model, direction)]['verdict']] for doi in doi_list], dtype=float)
- # Filtering
- trigger_rows = [r for r in raw if r['model'] in set(ALL_MODELS) and r['verdict'] in UNKNOWN]
- excluded_dois = sorted({r['doi'] for r in trigger_rows})
- retained_dois = [doi for doi in dois if doi not in set(excluded_dois)]
- n_total = len(dois)
- n_removed = len(excluded_dois)
- n_kept = len(retained_dois)
- flagged_by = defaultdict(set)
- for r in trigger_rows:
- flagged_by[r['doi']].add(r['model'])
- per_model_removed = {m: len({r['doi'] for r in trigger_rows if r['model'] == m}) for m in PRIMARY_MODELS}
- per_model_kept = {m: n_total - per_model_removed[m] for m in PRIMARY_MODELS}
- attribution = Counter()
- for doi, models in flagged_by.items():
- if len(models) == 1:
- attribution[next(iter(models))] += 1
- else:
- attribution['Multiple'] += 1
- print(f"Data loaded: {n_total} total DOIs, {n_kept} retained, {n_removed} removed")
- # %%
- # ============================================================
- # Plot 3: Per-model % kept vs % removed (Donut Charts)
- # ============================================================
- fig, axes = plt.subplots(1, len(PRIMARY_MODELS), figsize=(4.4 * len(PRIMARY_MODELS), 4.6))
- for ax, model in zip(axes, PRIMARY_MODELS):
- ax.pie([per_model_kept[model], per_model_removed[model]],
- colors=[KEPT_COLOR, REMOVED_COLOR],
- autopct=pct_count_autopct(n_total), startangle=90, counterclock=False,
- wedgeprops=dict(width=0.42, edgecolor='white'), pctdistance=0.75,
- textprops=dict(fontsize=8))
- ax.text(0, 0, DISPLAY[model], ha='center', va='center', fontsize=10, fontweight='bold')
- ax.set_title(f'{DISPLAY[model]}\nkept vs removed', fontsize=10)
- fig.subplots_adjust(bottom=0.2)
- handles = [plt.Rectangle((0, 0), 1, 1, color=KEPT_COLOR),
- plt.Rectangle((0, 0), 1, 1, color=REMOVED_COLOR)]
- fig.legend(handles, ['Kept (valid on both)', 'Removed (Unknown/Missing)'],
- ncol=2, frameon=False, loc='lower center', bbox_to_anchor=(0.5, 0.08))
- fig.suptitle(f'Per-model DOI pair kept vs removed (n={n_total} each)', fontsize=11)
- fig.text(0.5, 0.015, 'Per-model view: a pair is "removed" for that model if it returned Unknown/Missing on Right or Inverse (independent of the global filter).',
- ha='center', va='bottom', fontsize=6.5)
- plt.show()
- print("✅ Plot 3: Per-model Kept vs Removed Donuts")
- # %% [markdown]
- # ## 4c
- # %%
- # ============================================================
- # Plot 4: Ordinal distributions - Primary Models (Right)
- # ============================================================
- def ordinal_percent_matrix(models, direction, doi_list):
- matrix = np.zeros((len(models), len(VERDICTS)))
- for i, model in enumerate(models):
- counts = np.zeros(len(VERDICTS))
- for doi in doi_list:
- counts[SCORE[lookup_all[(doi, model, direction)]['verdict']]] += 1
- matrix[i] = counts / counts.sum() * 100
- return matrix
- models = PRIMARY_MODELS
- direction = 'Right'
- matrix = ordinal_percent_matrix(models, direction, retained_dois)
- fig, ax = plt.subplots(figsize=(9.2, 5.0))
- y = np.arange(len(models))
- left = np.zeros(len(models))
- for j, (label, color) in enumerate(zip(VERDICT_LABELS, VERDICT_COLORS)):
- ax.barh(y, matrix[:, j], left=left, label=label, color=color)
- for i in range(len(models)):
- value = matrix[i, j]
- if value >= 4.0:
- ax.text(left[i] + value / 2, i, f'{value:.0f}', ha='center', va='center', fontsize=8, color='white')
- left += matrix[:, j]
- ax.set_yticks(y, [DISPLAY[m] for m in models])
- ax.invert_yaxis()
- ax.set_xlim(0, 100)
- ax.set_xlabel('Verdict distribution (%)')
- ax.set_title(f'Normalized ordinal verdict distributions - {direction} hypotheses', pad=34)
- ax.legend(ncol=3, frameon=False, loc='lower center', bbox_to_anchor=(0.5, 1.02))
- ax.grid(axis='x', alpha=0.25)
- footer(fig, f'Right is experimentally supported; Inverse is the DOI-matched incorrect reverse hypothesis. Complete-case n={len(retained_dois)} pairs.')
- plt.show()
- print("✅ Plot 4: Ordinal Distributions - Primary Models (Right)")
- # %%
- # ============================================================
- # Plot 5: Ordinal distributions - Primary Models (Inverse)
- # ============================================================
- models = PRIMARY_MODELS
- direction = 'Inverse'
- matrix = ordinal_percent_matrix(models, direction, retained_dois)
- fig, ax = plt.subplots(figsize=(9.2, 5.0))
- y = np.arange(len(models))
- left = np.zeros(len(models))
- for j, (label, color) in enumerate(zip(VERDICT_LABELS, VERDICT_COLORS)):
- ax.barh(y, matrix[:, j], left=left, label=label, color=color)
- for i in range(len(models)):
- value = matrix[i, j]
- if value >= 4.0:
- ax.text(left[i] + value / 2, i, f'{value:.0f}', ha='center', va='center', fontsize=8, color='white')
- left += matrix[:, j]
- ax.set_yticks(y, [DISPLAY[m] for m in models])
- ax.invert_yaxis()
- ax.set_xlim(0, 100)
- ax.set_xlabel('Verdict distribution (%)')
- ax.set_title(f'Normalized ordinal verdict distributions - {direction} hypotheses', pad=34)
- ax.legend(ncol=3, frameon=False, loc='lower center', bbox_to_anchor=(0.5, 1.02))
- ax.grid(axis='x', alpha=0.25)
- footer(fig, f'Right is experimentally supported; Inverse is the DOI-matched incorrect reverse hypothesis. Complete-case n={len(retained_dois)} pairs.')
- plt.show()
- print("✅ Plot 5: Ordinal Distributions - Primary Models (Inverse)")
- # %% [markdown]
- # ## 4d
- # %%
- # ============================================================
- # Plot 7: Exact pair resolution (Lollipop)
- # ============================================================
- fig, ax = plt.subplots(figsize=(7.0, 4.4))
- y = np.arange(len(PRIMARY_MODELS))
- ax.hlines(y, 0, exact_rates, color='#9ecae1', linewidth=2.5, zorder=1)
- ax.scatter(exact_rates, y, color=BAR_COLOR, s=90, zorder=2)
- for yi, value in zip(y, exact_rates):
- ax.text(value + max(exact_rates + [0.01]) * 0.04 + 0.003, yi, f'{value:.1%}', va='center', fontsize=9)
- ax.set_yticks(y, [DISPLAY[m] for m in PRIMARY_MODELS])
- ax.invert_yaxis()
- ax.set_xlim(0, max(exact_rates + [0.05]) * 1.35 + 0.01)
- ax.set_xlabel('Exact DOI-pair resolution rate')
- ax.set_title('Right=Support and matched Inverse=Non-support')
- ax.grid(axis='x', alpha=0.25)
- footer(fig, 'Same metric as the bar chart, shown as a lollipop for comparison.')
- plt.show()
- print("✅ Plot 7: Exact Pair Resolution Lollipop")
- # %% [markdown]
- # ## 4f
- # %%
- # ============================================================================
- # PLOT 1: Trade-off: edge accuracy vs benchmark score
- # ============================================================================
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import matplotlib as mpl
- import re
- import os
- # Configuration
- PATH_BASE = "judge_results_base_resolved_common.csv"
- PATH_FT3 = "judge_results_3_2_resolved_common.csv"
- PATH_FT4 = "judge_results_4_2_resolved_common.csv"
- PATH_BENCH = "model-accuracy.csv"
- OUTDIR = "outputs"
- os.makedirs(OUTDIR, exist_ok=True)
- MODEL_ORDER = ["Base", "Model 1", "Model 2"]
- MODEL_COLORS = {"Base": "#2C5F9E", "Model 1": "#E8912D", "Model 2": "#3C8C42"}
- MODEL_MARKERS = {"Base": "o", "Model 1": "s", "Model 2": "^"}
- mpl.rcParams.update({
- "svg.fonttype": "none",
- "font.family": "sans-serif",
- "font.sans-serif": ["Arial", "Helvetica", "DejaVu Sans"],
- "font.weight": "bold",
- "axes.labelweight": "bold",
- "axes.titleweight": "bold",
- "axes.linewidth": 1.8,
- "axes.edgecolor": "black",
- "xtick.major.width": 1.6,
- "ytick.major.width": 1.6,
- "xtick.major.size": 6,
- "ytick.major.size": 6,
- "xtick.direction": "in",
- "ytick.direction": "in",
- "xtick.labelsize": 12,
- "ytick.labelsize": 12,
- "axes.labelsize": 13,
- "axes.titlesize": 14,
- "legend.fontsize": 11,
- "legend.edgecolor": "black",
- "legend.framealpha": 1.0,
- })
- def style_axes(ax):
- ax.tick_params(which="both", top=True, right=True, direction="in")
- for spine in ax.spines.values():
- spine.set_linewidth(1.8)
- spine.set_color("black")
- def resolve_prediction(row):
- pred = str(row.get("resolved_answer", "")).upper().strip()
- if pred in {"TRUE", "FALSE"}:
- return pred
- raw = str(row.get("judge_raw_response", "")).upper()
- has_true = bool(re.search(r"\bTRUE\b", raw))
- has_false = bool(re.search(r"\bFALSE\b", raw))
- if has_true and not has_false:
- return "TRUE"
- if has_false and not has_true:
- return "FALSE"
- return "OTHER"
- def load_judge_csv(path):
- df = pd.read_csv(path)
- if df["ground_truth"].dtype != bool:
- df["ground_truth"] = df["ground_truth"].astype(str).str.upper().map({"TRUE": True, "FALSE": False})
- df["prediction"] = df.apply(resolve_prediction, axis=1)
- df["answer_cat"] = df["prediction"].map({"TRUE": "Yes", "FALSE": "No", "OTHER": "Unresolved"})
- return df
- def compute_model_stats(df):
- n = len(df)
- committed = df[df["answer_cat"] != "Unresolved"]
- cols = ["Yes", "No", "Unresolved"]
- conf_counts = df.groupby(["ground_truth", "answer_cat"]).size().unstack(fill_value=0).reindex(index=[True, False], columns=cols, fill_value=0)
- conf_frac = conf_counts.div(conf_counts.sum(axis=1), axis=0)
- resolved_frac = len(committed) / n
- true_committed = committed[committed["ground_truth"] == True]
- false_committed = committed[committed["ground_truth"] == False]
- recall_true = (true_committed["answer_cat"] == "Yes").mean() * 100
- recall_false = (false_committed["answer_cat"] == "No").mean() * 100
- affirm_frac = (committed["answer_cat"] == "Yes").mean() * 100
- edge_accuracy = df["accuracy"].mean()
- prevalence = committed["ground_truth"].mean() * 100
- return {"n": n, "conf_counts": conf_counts, "conf_frac": conf_frac, "resolved_frac": resolved_frac,
- "recall_true": recall_true, "recall_false": recall_false, "affirm_frac": affirm_frac,
- "edge_accuracy": edge_accuracy, "prevalence": prevalence}
- dfs = {"Base": load_judge_csv(PATH_BASE), "Model 1": load_judge_csv(PATH_FT3), "Model 2": load_judge_csv(PATH_FT4)}
- stats = {name: compute_model_stats(df) for name, df in dfs.items()}
- bench_df = pd.read_csv(PATH_BENCH)
- bench_df["Model"] = bench_df["Model"].replace({"Finetuned 3": "Model 1", "Finetuned 4": "Model 2"})
- bench_cols = [c for c in bench_df.columns if c not in ("Model", "Node Loss (U/W)", "Edge Loss (U/W)", "Combined Loss (U/W)")]
- bench_df["avg_bench"] = bench_df[bench_cols].mean(axis=1)
- avg_bench = dict(zip(bench_df["Model"], bench_df["avg_bench"]))
- # Plot
- fig, ax = plt.subplots(figsize=(7.5, 6))
- label_offsets = {"Base": (0.012, 0.001), "Model 1": (-0.10, 0.006), "Model 2": (0.012, 0.003)}
- for name in MODEL_ORDER:
- x = stats[name]["edge_accuracy"]
- y = avg_bench[name]
- ax.scatter(x, y, s=260, marker=MODEL_MARKERS[name], color=MODEL_COLORS[name],
- edgecolor="black", linewidth=1.3, label=name, zorder=3)
- dx, dy = label_offsets[name]
- ax.annotate(name, (x, y), xytext=(x + dx, y + dy), fontsize=13, fontweight="bold", color=MODEL_COLORS[name])
- ax.set_xlabel("Edge accuracy (fraction correct, higher is better \u2192)")
- ax.set_ylabel("Average benchmark score (higher is better \u2191)")
- ax.set_title("Trade-off: edge accuracy vs benchmark score")
- ax.set_xlim(0.5, 1.05)
- style_axes(ax)
- legend = ax.legend(title="Model", loc="lower left", markerscale=0.8)
- legend.get_title().set_fontweight("bold")
- fig.tight_layout()
- fig.savefig(f"{OUTDIR}/fig1_tradeoff_accuracy_vs_benchmark.svg")
- plt.show()
- plt.close(fig)
- # %% [markdown]
- # ## 4g
- # %%
- # ============================================================================
- # PLOT 2: Confusion matrices: ground truth vs model's resolved answer
- # ============================================================================
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import matplotlib as mpl
- import re
- import os
- # Configuration
- PATH_BASE = "judge_results_base_resolved_common.csv"
- PATH_FT3 = "judge_results_3_2_resolved_common.csv"
- PATH_FT4 = "judge_results_4_2_resolved_common.csv"
- OUTDIR = "outputs"
- os.makedirs(OUTDIR, exist_ok=True)
- MODEL_ORDER = ["Base", "Model 1", "Model 2"]
- MODEL_COLORS = {"Base": "#2C5F9E", "Model 1": "#E8912D", "Model 2": "#3C8C42"}
- mpl.rcParams.update({
- "svg.fonttype": "none",
- "font.family": "sans-serif",
- "font.sans-serif": ["Arial", "Helvetica", "DejaVu Sans"],
- "font.weight": "bold",
- "axes.labelweight": "bold",
- "axes.titleweight": "bold",
- "axes.linewidth": 1.8,
- "axes.edgecolor": "black",
- "xtick.major.width": 1.6,
- "ytick.major.width": 1.6,
- "xtick.major.size": 6,
- "ytick.major.size": 6,
- "xtick.direction": "in",
- "ytick.direction": "in",
- "xtick.labelsize": 12,
- "ytick.labelsize": 12,
- "axes.labelsize": 13,
- "axes.titlesize": 14,
- "legend.fontsize": 11,
- "legend.edgecolor": "black",
- "legend.framealpha": 1.0,
- })
- def resolve_prediction(row):
- pred = str(row.get("resolved_answer", "")).upper().strip()
- if pred in {"TRUE", "FALSE"}:
- return pred
- raw = str(row.get("judge_raw_response", "")).upper()
- has_true = bool(re.search(r"\bTRUE\b", raw))
- has_false = bool(re.search(r"\bFALSE\b", raw))
- if has_true and not has_false:
- return "TRUE"
- if has_false and not has_true:
- return "FALSE"
- return "OTHER"
- def load_judge_csv(path):
- df = pd.read_csv(path)
- if df["ground_truth"].dtype != bool:
- df["ground_truth"] = df["ground_truth"].astype(str).str.upper().map({"TRUE": True, "FALSE": False})
- df["prediction"] = df.apply(resolve_prediction, axis=1)
- df["answer_cat"] = df["prediction"].map({"TRUE": "Yes", "FALSE": "No", "OTHER": "Unresolved"})
- return df
- def compute_model_stats(df):
- n = len(df)
- committed = df[df["answer_cat"] != "Unresolved"]
- cols = ["Yes", "No", "Unresolved"]
- conf_counts = df.groupby(["ground_truth", "answer_cat"]).size().unstack(fill_value=0).reindex(index=[True, False], columns=cols, fill_value=0)
- conf_frac = conf_counts.div(conf_counts.sum(axis=1), axis=0)
- resolved_frac = len(committed) / n
- true_committed = committed[committed["ground_truth"] == True]
- false_committed = committed[committed["ground_truth"] == False]
- recall_true = (true_committed["answer_cat"] == "Yes").mean() * 100
- recall_false = (false_committed["answer_cat"] == "No").mean() * 100
- affirm_frac = (committed["answer_cat"] == "Yes").mean() * 100
- edge_accuracy = df["accuracy"].mean()
- prevalence = committed["ground_truth"].mean() * 100
- return {"n": n, "conf_counts": conf_counts, "conf_frac": conf_frac, "resolved_frac": resolved_frac,
- "recall_true": recall_true, "recall_false": recall_false, "affirm_frac": affirm_frac,
- "edge_accuracy": edge_accuracy, "prevalence": prevalence}
- dfs = {"Base": load_judge_csv(PATH_BASE), "Model 1": load_judge_csv(PATH_FT3), "Model 2": load_judge_csv(PATH_FT4)}
- stats = {name: compute_model_stats(df) for name, df in dfs.items()}
- # Plot
- fig, axes = plt.subplots(1, 3, figsize=(15, 5.2))
- row_labels = ["Actually\nTrue", "Actually\nFalse"]
- col_labels = ["Yes\n(True)", "No\n(False)", "Unresolved"]
- im = None
- for ax, name in zip(axes, MODEL_ORDER):
- frac = stats[name]["conf_frac"].values
- counts = stats[name]["conf_counts"].values
- im = ax.imshow(frac, cmap="Blues", vmin=0, vmax=1, aspect="auto")
- for i in range(2):
- for j in range(3):
- val = frac[i, j]
- txt_color = "white" if val > 0.6 else "black"
- ax.text(j, i, f"{counts[i, j]:,}\n{val*100:.1f}%", ha="center", va="center",
- fontsize=12, fontweight="bold", color=txt_color)
- ax.set_xticks(range(3))
- ax.set_xticklabels(col_labels, fontsize=11)
- ax.set_yticks(range(2))
- ax.set_yticklabels(row_labels, fontsize=11)
- ax.set_title(name, color=MODEL_COLORS[name], fontsize=14, fontweight="bold")
- ax.set_xlabel("Model's answer")
- if name == "Base":
- ax.set_ylabel("Ground truth")
- for spine in ax.spines.values():
- spine.set_visible(True)
- spine.set_linewidth(1.8)
- spine.set_color("black")
- ax.tick_params(length=0)
- ax.set_xticks(np.arange(-0.5, 3, 1), minor=True)
- ax.set_yticks(np.arange(-0.5, 2, 1), minor=True)
- ax.grid(which="minor", color="white", linewidth=2)
- ax.tick_params(which="minor", length=0)
- fig.suptitle("Confusion matrices: ground truth vs model's resolved answer", fontsize=15, fontweight="bold")
- cbar = fig.colorbar(im, ax=axes, fraction=0.025, pad=0.02)
- cbar.set_label("Row-normalised fraction", fontsize=11, fontweight="bold")
- fig.savefig(f"{OUTDIR}/fig2_confusion_matrices.svg", bbox_inches="tight")
- plt.show()
- plt.close(fig)
- # %% [markdown]
- # ## 4h
- # %%
- # ============================================================
- # Plot 8: Ordinal distributions - Alternative Models (Right)
- # ============================================================
- models = ALT_MODELS
- direction = 'Right'
- matrix = ordinal_percent_matrix(models, direction, retained_dois)
- fig, ax = plt.subplots(figsize=(9.2, 5.0))
- y = np.arange(len(models))
- left = np.zeros(len(models))
- for j, (label, color) in enumerate(zip(VERDICT_LABELS, VERDICT_COLORS)):
- ax.barh(y, matrix[:, j], left=left, label=label, color=color)
- for i in range(len(models)):
- value = matrix[i, j]
- if value >= 4.0:
- ax.text(left[i] + value / 2, i, f'{value:.0f}', ha='center', va='center', fontsize=8, color='white')
- left += matrix[:, j]
- ax.set_yticks(y, [DISPLAY[m] for m in models])
- ax.invert_yaxis()
- ax.set_xlim(0, 100)
- ax.set_xlabel('Verdict distribution (%)')
- ax.set_title(f'Normalized ordinal verdict distributions - {direction} hypotheses', pad=34)
- ax.legend(ncol=3, frameon=False, loc='lower center', bbox_to_anchor=(0.5, 1.02))
- ax.grid(axis='x', alpha=0.25)
- footer(fig, f'Right is experimentally supported; Inverse is the DOI-matched incorrect reverse hypothesis. Complete-case n={len(retained_dois)} pairs.')
- plt.show()
- print("✅ Plot 8: Ordinal Distributions - Alternative Models (Right)")
- # %%
- # ============================================================
- # Plot 9: Ordinal distributions - Alternative Models (Inverse)
- # ============================================================
- models = ALT_MODELS
- direction = 'Inverse'
- matrix = ordinal_percent_matrix(models, direction, retained_dois)
- fig, ax = plt.subplots(figsize=(9.2, 5.0))
- y = np.arange(len(models))
- left = np.zeros(len(models))
- for j, (label, color) in enumerate(zip(VERDICT_LABELS, VERDICT_COLORS)):
- ax.barh(y, matrix[:, j], left=left, label=label, color=color)
- for i in range(len(models)):
- value = matrix[i, j]
- if value >= 4.0:
- ax.text(left[i] + value / 2, i, f'{value:.0f}', ha='center', va='center', fontsize=8, color='white')
- left += matrix[:, j]
- ax.set_yticks(y, [DISPLAY[m] for m in models])
- ax.invert_yaxis()
- ax.set_xlim(0, 100)
- ax.set_xlabel('Verdict distribution (%)')
- ax.set_title(f'Normalized ordinal verdict distributions - {direction} hypotheses', pad=34)
- ax.legend(ncol=3, frameon=False, loc='lower center', bbox_to_anchor=(0.5, 1.02))
- ax.grid(axis='x', alpha=0.25)
- footer(fig, f'Right is experimentally supported; Inverse is the DOI-matched incorrect reverse hypothesis. Complete-case n={len(retained_dois)} pairs.')
- plt.show()
- print("✅ Plot 9: Ordinal Distributions - Alternative Models (Inverse)")
- # %% [markdown]
- # ## 4i
- # %%
- """
- Single-plot version: verdicts by level, drawn as a lollipop (dot) plot.
- """
- %matplotlib inline
- import ast
- import os
- # Remove matplotlib.use("Agg") - let it use default backend
- import matplotlib.pyplot as plt
- import numpy as np
- import pandas as pd
- # Fix for __file__ not being defined in interactive environments
- try:
- HERE = os.path.dirname(os.path.abspath(__file__))
- except NameError:
- HERE = os.getcwd()
- OUTDIR = os.path.join(HERE, "verdict_analysis")
- # Updated paths for BioChatter and Escargot
- BASE_PATH = "/storage/Arushi/090526_EvoAge/kg_formation/all_figures/FIG_4"
- BIOCHATTER_CSV = os.path.join(BASE_PATH, "biochatter_hypothesis_results_medgemma.csv")
- ESCARGOT_CSV = os.path.join(BASE_PATH, "escargot_hypothesis_results_medgemma.csv")
- EVOAGE_CSV = ("/storage/Arushi/090526_EvoAge/multiagent_hypo/"
- "evoage_100_right_inverse/Right_hypothesis_Extract_entities/"
- "hypothesis_pipeline_results_final_with_percentile_buckets_full.csv")
- LADDER = ["no_support", "weak_support", "partial_support", "support", "strong_support"]
- LABEL = ["No\nsupport", "Weak\nsupport", "Partial\nsupport", "Support", "Strong\nsupport"]
- RANK = {v: i for i, v in enumerate(LADDER)}
- SYSTEMS = ["EvoAge", "Escargot", "BioChatter"]
- COLOR = {"EvoAge": "#2a78d6", "Escargot": "#eb6834", "BioChatter": "#1baf7a"}
- INK, INK_SOFT, GRID, SURFACE = "#0b0b0b", "#52514e", "#e3e2de", "#fcfcfb"
- def parse_evoage_verdict(value):
- try:
- return (ast.literal_eval(value) or {}).get("verdict")
- except (ValueError, SyntaxError, TypeError):
- return None
- def load() -> pd.DataFrame:
- def key(series):
- return series.astype(str).str.strip().str.lower().str.rstrip("/")
- def clean(series):
- cleaned = series.astype(str).str.strip().str.lower()
- return cleaned.where(cleaned.isin(LADDER))
- for csv_path in [EVOAGE_CSV, ESCARGOT_CSV, BIOCHATTER_CSV]:
- if not os.path.exists(csv_path):
- print(f"Warning: File not found: {csv_path}")
- evo = pd.read_csv(EVOAGE_CSV)
- esc = pd.read_csv(ESCARGOT_CSV)
- bio = pd.read_csv(BIOCHATTER_CSV)
- evo = evo.assign(EvoAge=evo["EvoAge_hypothesis_response"].apply(parse_evoage_verdict),
- _key=key(evo["DOI"]))
- esc = esc.assign(Escargot=esc["Medgemma_verdict"], _key=key(esc["DOI"]))
- bio = bio.assign(BioChatter=bio["Medgemma_verdict"], _key=key(bio["DOI"]))
- merged = (evo[["_key", "Title", "DOI", "Right Hypothesis", "EvoAge"]]
- .merge(esc[["_key", "Escargot"]], on="_key", how="inner", validate="1:1")
- .merge(bio[["_key", "BioChatter"]], on="_key", how="inner", validate="1:1"))
- for system in SYSTEMS:
- merged[system] = clean(merged[system])
- return merged.drop(columns="_key")
- def plot_lollipop(frame):
- fig, ax = plt.subplots(figsize=(7.6, 5.0), facecolor=SURFACE)
- ax.set_facecolor(SURFACE)
- counts = {s: frame[s].value_counts().reindex(LADDER).fillna(0).astype(int)
- for s in SYSTEMS}
- y_base = np.arange(len(LADDER))[::-1]
- offsets = {"EvoAge": 0.23, "Escargot": 0.0, "BioChatter": -0.23}
- for system in SYSTEMS:
- y = y_base + offsets[system]
- values = counts[system].values
- ax.hlines(y, 0, values, color=COLOR[system], linewidth=1.8, alpha=0.55)
- ax.plot(values, y, "o", markersize=8, color=COLOR[system],
- markeredgecolor=SURFACE, markeredgewidth=1.4,
- label=system, linestyle="none", zorder=3)
- for yi, value in zip(y, values):
- if value > 0:
- ax.text(value + 2.4, yi, str(value), va="center",
- fontsize=8.5, color=INK_SOFT)
- for side in ("top", "right"):
- ax.spines[side].set_visible(False)
- for side in ("left", "bottom"):
- ax.spines[side].set_color(GRID)
- ax.tick_params(colors=INK_SOFT, labelsize=9, length=3)
- ax.set_axisbelow(True)
- ax.xaxis.grid(True, color=GRID, linewidth=0.7)
- ax.set_yticks(y_base)
- ax.set_yticklabels([l.replace("\n", " ") for l in LABEL], fontsize=10, color=INK)
- ax.set_ylim(-0.6, len(LADDER) - 0.4)
- ax.set_xlim(0, max(108, len(frame) * 1.08))
- ax.set_xlabel("Hypotheses (n)", fontsize=10, color=INK_SOFT)
- ax.set_title(f"Verdicts by level across {len(frame)} biological hypotheses",
- loc="left", fontsize=12, color=INK, pad=12, fontweight="bold")
- ax.legend(frameon=False, fontsize=10, loc="lower right", labelcolor=INK_SOFT)
- fig.tight_layout()
- return fig
- def main() -> None:
- os.makedirs(OUTDIR, exist_ok=True)
- print(f"BioChatter file: {BIOCHATTER_CSV}")
- print(f"Escargot file: {ESCARGOT_CSV}")
- print(f"EvoAge file: {EVOAGE_CSV}")
- frame = load()
- print(f"joined {len(frame)} hypotheses on DOI")
- for system in SYSTEMS:
- n = frame[system].isin(LADDER[1:]).sum()
- print(f" {system:<11} any support: {n}/{len(frame)} ({100*n/len(frame):.0f}%)")
- fig = plot_lollipop(frame)
- svg = os.path.join(OUTDIR, "figure4_verdict_lollipop.svg")
- png = os.path.join(OUTDIR, "figure4_verdict_lollipop.png")
- fig.savefig(svg, format="svg", facecolor=SURFACE, bbox_inches="tight")
- fig.savefig(png, dpi=200, facecolor=SURFACE, bbox_inches="tight")
- plt.show() # This will display the plot
- plt.close(fig)
- print(f"\nwrote {svg}\nwrote {png}")
- if __name__ == "__main__":
- main()
- # %% [markdown]
- # ## Main Figure 5
- # %% [markdown]
- # ## 5b
- # %%
- # This magic command makes plots appear in the notebook
- %matplotlib inline
- """
- log2fc_coherence_plot.py
- Cross-condition coherence plot using log2 fold change (log2FC) instead of
- difference scores.
- For each gene:
- log2FC_c1 = log2(Mean_KO_ratio_c1 / Mean_Control_ratio_c1) # Heat shock 37C vs 30C
- log2FC_c2 = log2(Mean_KO_ratio_c2 / Mean_Control_ratio_c2) # Thermo pulsing vs 30C
- Color = significance status (BH-FDR q <= 0.05), NOT EvoAge verdict:
- grey = not significant
- red = heat shock significant only
- blue = thermo pulsing significant only
- green = significant in both conditions
- Spearman rho computed across all 976 genes (no filtering).
- INPUT : /storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/03_log2fc_coherence/input_EvoAge_log2FC_merged_data.csv
- OUTPUT: coherence_log2fc.png / .svg in the same directory
- """
- import warnings
- warnings.filterwarnings('ignore')
- import matplotlib.pyplot as plt
- import numpy as np
- import pandas as pd
- from matplotlib.lines import Line2D
- from scipy import stats
- import os
- from pathlib import Path
- # ── Configuration ──────────────────────────────────────────────────────
- # Input file path
- INPUT_FILE = '/storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/03_log2fc_coherence/input_EvoAge_log2FC_merged_data.csv'
- # Output directory (same as input file location)
- OUTPUT_DIR = os.path.dirname(INPUT_FILE)
- OUTPUT_BASE = 'coherence_log2fc'
- # ── Gene name mapping (verified from EvoAge_complete_merged_data.csv) ──
- GENE_NAMES = {
- "YGR036C": "CAX4", "YDL219W": "DTD1", "YDR098C": "GRX3", "YGR155W": "CYS4",
- "YDR226W": "ADK1", "YBR035C": "PDX3", "YBL099W": "ATP1", "YGL012W": "ERG4",
- "YGR087C": "PDC6", "YOR316C": "COT1",
- }
- # Label offsets to prevent text overlap
- LABEL_OFFSETS = {
- "CAX4": (8, 8), "DTD1": (8, -10), "GRX3": (-10, 8), "CYS4": (8, -10), "ADK1": (8, 8),
- "PDX3": (8, 8), "ATP1": (8, -12), "ERG4": (-12, 8), "PDC6": (8, -10), "COT1": (-12, -12),
- }
- # Significance-based colors
- COLORS = {
- "not_sig": "#B0B0B0", # grey
- "heat_only": "#C0392B", # red
- "pulse_only": "#2874A6", # blue
- "both": "#2E8B57", # green
- }
- def load_and_compute(csv_path):
- """Load merged data and compute log2FC if not already present."""
- print(f"Loading data from: {csv_path}")
- df = pd.read_csv(csv_path)
- print(f"Loaded {len(df)} genes")
- # Compute log2FC if the columns don't already exist
- if "log2FC_c1" not in df.columns:
- df["log2FC_c1"] = np.log2(df["Mean_KO_ratio_c1"] / df["Mean_Control_ratio_c1"])
- if "log2FC_c2" not in df.columns:
- df["log2FC_c2"] = np.log2(df["Mean_KO_ratio_c2"] / df["Mean_Control_ratio_c2"])
- # Significance from BH-FDR q-values (these are NOT changed by log2FC)
- df["sig37"] = df["q_value_c1"] <= 0.05
- df["sigPul"] = df["q_value_c2"] <= 0.05
- df["sig_any"] = df["sig37"] | df["sigPul"]
- # Significance category for coloring
- df["color"] = COLORS["not_sig"]
- df.loc[df["sig37"] & ~df["sigPul"], "color"] = COLORS["heat_only"]
- df.loc[df["sigPul"] & ~df["sig37"], "color"] = COLORS["pulse_only"]
- df.loc[df["sig37"] & df["sigPul"], "color"] = COLORS["both"]
- return df.dropna(subset=["log2FC_c1", "log2FC_c2"])
- def make_plot(df, output_dir, output_base, show_plot=True):
- """Create the log2FC cross-condition coherence plot."""
- x = df["log2FC_c1"]
- y = df["log2FC_c2"]
- sig_any = df["sig_any"]
- fig, ax = plt.subplots(figsize=(7.5, 7))
- # All genes (grey majority, colored minority)
- ax.scatter(x, y, c=df["color"], s=30, alpha=0.55,
- edgecolors="none", rasterized=True, zorder=2)
- # Significant genes on top with black outline
- ax.scatter(x[sig_any], y[sig_any], s=90,
- facecolors=df.loc[sig_any, "color"],
- edgecolors="black", linewidths=1.0, zorder=5)
- # Label the 10 significant genes
- for gid, name in GENE_NAMES.items():
- row = df[df["Gene"] == gid]
- if len(row):
- r = row.iloc[0]
- dx, dy = LABEL_OFFSETS.get(name, (8, 8))
- ax.annotate(
- name, (r["log2FC_c1"], r["log2FC_c2"]),
- fontsize=8.5, fontweight="bold",
- ha="left" if dx >= 0 else "right",
- va="bottom" if dy >= 0 else "top",
- xytext=(dx, dy), textcoords="offset points",
- bbox=dict(boxstyle="round,pad=0.15", facecolor="white",
- edgecolor="none", alpha=0.85),
- zorder=6,
- )
- # y=x diagonal + zero lines
- lo, hi = -1.3, 2.0
- ax.plot([lo, hi], [lo, hi], "k--", alpha=0.35, lw=0.8, zorder=1)
- ax.axhline(0, color="#C3C6CA", lw=0.7, zorder=1)
- ax.axvline(0, color="#C3C6CA", lw=0.7, zorder=1)
- # Spearman (all genes, no filtering)
- rho, p = stats.spearmanr(x, y)
- ax.text(
- 0.03, 0.97,
- f"Spearman ρ = {rho:.2f}\nP = {p:.1e}\nn = {len(df)}",
- transform=ax.transAxes, va="top", ha="left", fontsize=10,
- bbox=dict(boxstyle="round", facecolor="white", alpha=0.9), zorder=7,
- )
- ax.set_xlim(lo, hi)
- ax.set_ylim(lo, hi)
- ax.set_xlabel("Heat shock: log₂(KO/control)", fontsize=12)
- ax.set_ylabel("Thermo pulsing: log₂(KO/control)", fontsize=12)
- ax.set_title("Cross-condition coherence (log₂FC, n=976)", fontsize=13, fontweight="bold")
- # Legend
- legend_elems = [
- Line2D([0], [0], marker="o", color="w", markerfacecolor=COLORS["not_sig"], markersize=9, label="Not significant"),
- Line2D([0], [0], marker="o", color="w", markerfacecolor=COLORS["heat_only"], markersize=9, label="Heat shock significant"),
- Line2D([0], [0], marker="o", color="w", markerfacecolor=COLORS["pulse_only"], markersize=9, label="Thermo pulsing significant"),
- Line2D([0], [0], marker="o", color="w", markerfacecolor=COLORS["both"], markersize=9, label="Both significant"),
- ]
- ax.legend(handles=legend_elems, loc="lower right", fontsize=8,
- frameon=True, edgecolor="grey")
- ax.spines["top"].set_visible(False)
- ax.spines["right"].set_visible(False)
- plt.tight_layout()
- # Show plot in notebook
- if show_plot:
- plt.show()
- # Save figures
- for ext in ['png', 'svg']:
- output_path = os.path.join(output_dir, f"{output_base}.{ext}")
- plt.savefig(output_path, dpi=200, bbox_inches="tight")
- print(f"✅ Saved: {output_path}")
- plt.close()
- print(f"\n📊 Statistics:")
- print(f"Spearman ρ = {rho:.5f}, P = {p:.2e}")
- print(f"Genes: n={len(df)}")
- print(f" Not significant: {(~sig_any).sum()}")
- print(f" Heat shock only: {(df['sig37'] & ~df['sigPul']).sum()}")
- print(f" Thermo pulsing only: {(df['sigPul'] & ~df['sig37']).sum()}")
- print(f" Both significant: {(df['sig37'] & df['sigPul']).sum()}")
- # ── Main execution ──────────────────────────────────────────────────────
- # Load data
- df = load_and_compute(INPUT_FILE)
- # Generate plot (will show in notebook and save to files)
- make_plot(df, OUTPUT_DIR, OUTPUT_BASE, show_plot=True)
- print(f"\n✅ All done! Plots saved in: {OUTPUT_DIR}")
- # %% [markdown]
- # ## 5c
- # %%
- # ============================================
- # PLOT 1: Heat Shock (37 °C) vs 30 °C
- # ============================================
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- from matplotlib.lines import Line2D
- import matplotlib.ticker as ticker
- import os
- import warnings
- # Suppress font warnings
- warnings.filterwarnings('ignore', category=UserWarning, module='matplotlib')
- # Set font to a standard one that exists in most systems
- plt.rcParams['font.family'] = 'sans-serif'
- plt.rcParams['font.sans-serif'] = ['Arial', 'DejaVu Sans', 'Helvetica', 'sans-serif']
- # Enable inline plotting in Jupyter
- %matplotlib inline
- # Define colors and labels
- VERDICT_COLORS = {
- 'support': '#6FAE84',
- 'weak_support': '#7E7AAE',
- 'partial_support': '#E0D55A',
- 'no_support': '#D9754A',
- }
- VERDICT_LABELS = {
- 'support': 'Support',
- 'weak_support': 'Weak support',
- 'partial_support': 'Partial support',
- 'no_support': 'No support',
- }
- def to_bool(s):
- return str(s).strip().lower() in {'true','1','yes'}
- def prep(df):
- out = df.copy()
- out['Sig_37C'] = out['Sig_37C'].map(to_bool)
- out['Sig_Pulser'] = out['Sig_Pulser'].map(to_bool)
- out['neglog10_q_37C'] = -np.log10(pd.to_numeric(out['q_value_c1'], errors='coerce').clip(lower=1e-300))
- out['neglog10_q_Pulser'] = -np.log10(pd.to_numeric(out['q_value_c2'], errors='coerce').clip(lower=1e-300))
- out['Score_c1'] = pd.to_numeric(out['Score_c1'], errors='coerce')
- out['Score_c2'] = pd.to_numeric(out['Score_c2'], errors='coerce')
- # Use standard names like ERG4, ATP1; fall back to systematic names if needed
- out['label_name'] = out['Standard_Name'].fillna('').astype(str)
- out.loc[out['label_name'].str.strip()=='', 'label_name'] = out.loc[out['label_name'].str.strip()=='', 'Gene']
- return out
- def style_axis_heatshock(ax):
- """Style the axis for heat shock plot"""
- ax.set_title('Heat shock (37 °C) vs 30 °C', fontsize=12, pad=8)
- ax.text(0.5, 0.985, 'n = 976 genes; all FDR values shown', transform=ax.transAxes,
- ha='center', va='top', fontsize=8)
- ax.set_xlabel('Thermal-protection score', fontsize=11)
- ax.set_ylabel(r'-log$_{10}$(BH FDR)', fontsize=11)
- ax.axvline(0, color='#D0D0D0', lw=1.2, zorder=0)
- qline = -np.log10(0.05)
- ax.axhline(qline, color='#9A9A9A', lw=1.6, ls=(0,(4,3)), zorder=0)
- ax.text(0.985, qline+0.05, 'dashed line: q = 0.05', transform=ax.get_yaxis_transform(),
- ha='right', va='bottom', fontsize=8, color='#555555')
- ax.set_xlim(-0.7, 2.6)
- ax.set_xticks([-0.5, 0, 0.5, 1, 1.5, 2, 2.5])
- ax.set_ylim(0, 110)
- ax.set_yscale('symlog', linthresh=1.0, linscale=0.85, base=10)
- ax.set_yticks([0, 1, 2, 5, 10, 20, 50, 100])
- ax.get_yaxis().set_major_formatter(ticker.ScalarFormatter())
- ax.tick_params(width=1.0, length=4, labelsize=9)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- for s in ['left','bottom']:
- ax.spines[s].set_linewidth(1.1)
- def make_handles():
- handles = [Line2D([0],[0], marker='o', linestyle='None', markersize=4,
- markerfacecolor='#C9C9C9', markeredgecolor='none', label='Not significant')]
- for key in ['support','weak_support','partial_support','no_support']:
- handles.append(Line2D([0],[0], marker='o', linestyle='None', markersize=5,
- markerfacecolor=VERDICT_COLORS[key], markeredgecolor='none',
- label=VERDICT_LABELS[key]))
- return handles
- def label_points_heatshock(ax, subdf):
- """Label points for heat shock plot"""
- offsets = {
- 'YGR036C': (-0.04, 0.0, 'right'), # CAX4
- 'YDL219W': (0.04, 0.0, 'left'), # DTD1
- 'YDR098C': (0.04, 0.0, 'left'), # GRX3
- 'YGR155W': (0.04, 0.0, 'left'), # CYS4
- 'YDR226W': (0.04, -0.08, 'left'), # ADK1
- 'YBR035C': (0.04, -0.06, 'left'), # PDX3
- 'YBL099W': (0.04, 0.0, 'left'), # ATP1
- 'YGL012W': (0.04, 0.02, 'left'), # ERG4
- 'YGR087C': (0.04, -0.06, 'left'), # PDC6
- 'YOR316C': (-0.05, 0.02, 'right'), # COT1
- }
- for _, r in subdf.iterrows():
- dx, dy, ha = offsets.get(r['Gene'], (0.04, 0.0, 'left'))
- ax.text(r['Score_c1']+dx, r['neglog10_q_37C']+dy, r['label_name'],
- fontsize=8.8, ha=ha, va='center', color='#333333')
- # ============================================
- # LOAD DATA AND CREATE HEAT SHOCK PLOT
- # ============================================
- # Update this path to your CSV file
- input_file = 'EvoAge_effect_FDR_input.csv'
- if not os.path.exists(input_file):
- print(f"⚠️ Error: File '{input_file}' not found!")
- print(f"Current directory: {os.getcwd()}")
- else:
- # Load and prepare data
- print(f"✅ Loading data from: {input_file}")
- df = prep(pd.read_csv(input_file))
- print(f"✅ Data loaded. Shape: {df.shape}")
- print(f"✅ Significant genes at 37°C: {df['Sig_37C'].sum()}")
- # Create the plot
- fig, ax = plt.subplots(figsize=(4.1,4.0), constrained_layout=True)
- # Style the axis
- style_axis_heatshock(ax)
- # Plot all points in grey
- ax.scatter(df['Score_c1'], df['neglog10_q_37C'], s=10, color='#C9C9C9',
- alpha=0.85, edgecolors='none', zorder=1)
- # Plot significant points with colors based on verdict
- sig = df[df['Sig_37C']].copy()
- for key in ['support','weak_support','partial_support','no_support']:
- ss = sig[sig['verdict'] == key]
- if len(ss):
- ax.scatter(ss['Score_c1'], ss['neglog10_q_37C'], s=24,
- color=VERDICT_COLORS[key], edgecolors='none', zorder=3)
- # Add labels
- label_points_heatshock(ax, sig)
- # Add legend
- ax.legend(handles=make_handles(), title='EvoAge verdict', loc='lower right',
- frameon=False, fontsize=8, title_fontsize=8.5,
- handletextpad=0.4, borderpad=0.2, labelspacing=0.3)
- # Add panel label
- ax.text(-0.13, 1.04, 'c', transform=ax.transAxes, fontsize=16, fontweight='bold')
- # Display the plot
- plt.show()
- # Optional: Save the figure (uncomment to use)
- # fig.savefig('heatshock_panel.png', bbox_inches='tight', dpi=300)
- # fig.savefig('heatshock_panel.pdf', bbox_inches='tight')
- # print("✅ Figure saved!")
- plt.close(fig)
- # %% [markdown]
- # ## 5d
- # %%
- # ============================================
- # PLOT 2: Thermo Pulsing vs 30 °C
- # ============================================
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- from matplotlib.lines import Line2D
- import matplotlib.ticker as ticker
- import os
- import warnings
- # Suppress font warnings
- warnings.filterwarnings('ignore', category=UserWarning, module='matplotlib')
- # Set font to a standard one that exists in most systems
- plt.rcParams['font.family'] = 'sans-serif'
- plt.rcParams['font.sans-serif'] = ['Arial', 'DejaVu Sans', 'Helvetica', 'sans-serif']
- # Enable inline plotting in Jupyter
- %matplotlib inline
- # Define colors and labels
- VERDICT_COLORS = {
- 'support': '#6FAE84',
- 'weak_support': '#7E7AAE',
- 'partial_support': '#E0D55A',
- 'no_support': '#D9754A',
- }
- VERDICT_LABELS = {
- 'support': 'Support',
- 'weak_support': 'Weak support',
- 'partial_support': 'Partial support',
- 'no_support': 'No support',
- }
- def to_bool(s):
- return str(s).strip().lower() in {'true','1','yes'}
- def prep(df):
- out = df.copy()
- out['Sig_37C'] = out['Sig_37C'].map(to_bool)
- out['Sig_Pulser'] = out['Sig_Pulser'].map(to_bool)
- out['neglog10_q_37C'] = -np.log10(pd.to_numeric(out['q_value_c1'], errors='coerce').clip(lower=1e-300))
- out['neglog10_q_Pulser'] = -np.log10(pd.to_numeric(out['q_value_c2'], errors='coerce').clip(lower=1e-300))
- out['Score_c1'] = pd.to_numeric(out['Score_c1'], errors='coerce')
- out['Score_c2'] = pd.to_numeric(out['Score_c2'], errors='coerce')
- # Use standard names like ERG4, ATP1; fall back to systematic names if needed
- out['label_name'] = out['Standard_Name'].fillna('').astype(str)
- out.loc[out['label_name'].str.strip()=='', 'label_name'] = out.loc[out['label_name'].str.strip()=='', 'Gene']
- return out
- def style_axis_thermopulsing(ax):
- """Style the axis for thermo pulsing plot"""
- ax.set_title('Thermo pulsing vs 30 °C', fontsize=12, pad=8)
- ax.text(0.5, 0.985, 'n = 976 genes; all FDR values shown', transform=ax.transAxes,
- ha='center', va='top', fontsize=8)
- ax.set_xlabel('Thermal-protection score', fontsize=11)
- ax.set_ylabel(r'-log$_{10}$(BH FDR)', fontsize=11)
- ax.axvline(0, color='#D0D0D0', lw=1.2, zorder=0)
- qline = -np.log10(0.05)
- ax.axhline(qline, color='#9A9A9A', lw=1.6, ls=(0,(4,3)), zorder=0)
- ax.text(0.985, qline+0.05, 'dashed line: q = 0.05', transform=ax.get_yaxis_transform(),
- ha='right', va='bottom', fontsize=8, color='#555555')
- ax.set_xlim(-0.7, 2.6)
- ax.set_xticks([-0.5, 0, 0.5, 1, 1.5, 2, 2.5])
- ax.set_ylim(0, 110)
- ax.set_yscale('symlog', linthresh=1.0, linscale=0.85, base=10)
- ax.set_yticks([0, 1, 2, 5, 10, 20, 50, 100])
- ax.get_yaxis().set_major_formatter(ticker.ScalarFormatter())
- ax.tick_params(width=1.0, length=4, labelsize=9)
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- for s in ['left','bottom']:
- ax.spines[s].set_linewidth(1.1)
- def make_handles():
- handles = [Line2D([0],[0], marker='o', linestyle='None', markersize=4,
- markerfacecolor='#C9C9C9', markeredgecolor='none', label='Not significant')]
- for key in ['support','weak_support','partial_support','no_support']:
- handles.append(Line2D([0],[0], marker='o', linestyle='None', markersize=5,
- markerfacecolor=VERDICT_COLORS[key], markeredgecolor='none',
- label=VERDICT_LABELS[key]))
- return handles
- def label_points_thermopulsing(ax, subdf):
- """Label points for thermo pulsing plot"""
- offsets = {
- 'YGR036C': (-0.04, 0.0, 'right'), # CAX4
- 'YDL219W': (0.04, 0.0, 'left'), # DTD1
- 'YDR098C': (0.04, 0.0, 'left'), # GRX3
- 'YGR155W': (0.04, 0.0, 'left'), # CYS4
- 'YDR226W': (0.04, -0.08, 'left'), # ADK1
- 'YBR035C': (0.04, -0.06, 'left'), # PDX3
- 'YBL099W': (0.04, 0.0, 'left'), # ATP1
- 'YGL012W': (0.04, 0.02, 'left'), # ERG4
- 'YGR087C': (0.04, -0.06, 'left'), # PDC6
- 'YOR316C': (-0.05, 0.02, 'right'), # COT1
- }
- for _, r in subdf.iterrows():
- dx, dy, ha = offsets.get(r['Gene'], (0.04, 0.0, 'left'))
- ax.text(r['Score_c2']+dx, r['neglog10_q_Pulser']+dy, r['label_name'],
- fontsize=8.8, ha=ha, va='center', color='#333333')
- # ============================================
- # LOAD DATA AND CREATE THERMO PULSING PLOT
- # ============================================
- # Update this path to your CSV file
- input_file = 'EvoAge_effect_FDR_input.csv'
- if not os.path.exists(input_file):
- print(f"⚠️ Error: File '{input_file}' not found!")
- print(f"Current directory: {os.getcwd()}")
- else:
- # Load and prepare data
- print(f"✅ Loading data from: {input_file}")
- df = prep(pd.read_csv(input_file))
- print(f"✅ Data loaded. Shape: {df.shape}")
- print(f"✅ Significant genes in pulser: {df['Sig_Pulser'].sum()}")
- # Create the plot
- fig, ax = plt.subplots(figsize=(4.1,4.0), constrained_layout=True)
- # Style the axis
- style_axis_thermopulsing(ax)
- # Plot all points in grey
- ax.scatter(df['Score_c2'], df['neglog10_q_Pulser'], s=10, color='#C9C9C9',
- alpha=0.85, edgecolors='none', zorder=1)
- # Plot significant points with colors based on verdict
- sig = df[df['Sig_Pulser']].copy()
- for key in ['support','weak_support','partial_support','no_support']:
- ss = sig[sig['verdict'] == key]
- if len(ss):
- ax.scatter(ss['Score_c2'], ss['neglog10_q_Pulser'], s=24,
- color=VERDICT_COLORS[key], edgecolors='none', zorder=3)
- # Add labels
- label_points_thermopulsing(ax, sig)
- # Add legend
- ax.legend(handles=make_handles(), title='EvoAge verdict', loc='lower right',
- frameon=False, fontsize=8, title_fontsize=8.5,
- handletextpad=0.4, borderpad=0.2, labelspacing=0.3)
- # Add panel label
- ax.text(-0.13, 1.04, 'd', transform=ax.transAxes, fontsize=16, fontweight='bold')
- # Display the plot
- plt.show()
- # Optional: Save the figure (uncomment to use)
- # fig.savefig('thermopulsing_panel.png', bbox_inches='tight', dpi=300)
- # fig.savefig('thermopulsing_panel.pdf', bbox_inches='tight')
- # print("✅ Figure saved!")
- plt.close(fig)
- # %% [markdown]
- # ## 5e
- # %%
- # This magic command makes plots appear in the notebook
- %matplotlib inline
- """
- EvoAge Overlap Venn Diagram
- ---------------------------
- Creates a Venn diagram showing overlap of significant genes between conditions.
- INPUT : /storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/overlap/EvoAge_complete_merged_data.csv
- OUTPUT: 03_overlap_venn.svg / .png in the same directory
- """
- import warnings
- warnings.filterwarnings('ignore')
- from pathlib import Path
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- from matplotlib.patches import Circle
- # ── Configuration ──────────────────────────────────────────────────────
- # Input directory and file
- INPUT_DIR = '/storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/overlap'
- INPUT_FILE = Path(INPUT_DIR) / 'EvoAge_complete_merged_data.csv'
- OUTPUT_DIR = INPUT_DIR
- # ── Constants (from common.py) ──────────────────────────────────────
- VERDICT_MAP = {
- 'support': 'Support',
- 'weak_support': 'Weak support',
- 'partial_support': 'Partial support',
- 'no_support': 'No support',
- }
- VERDICT_ORDER = ['Support', 'Weak support', 'Partial support', 'No support', 'Unresolved']
- COLORS = {
- 'Support': '#5B4B8A',
- 'Weak support': '#6FAE9B',
- 'Partial support': '#D8B65C',
- 'No support': '#C97B63',
- 'Unresolved': '#CFCFD4',
- '37C': '#B66A5B',
- 'Pulser': '#6E9D8C',
- 'neutral': '#C9CDD2',
- }
- def configure():
- plt.rcParams.update({
- 'font.family': 'DejaVu Sans',
- 'font.size': 9,
- 'axes.titlesize': 11,
- 'axes.labelsize': 10,
- 'xtick.labelsize': 8.5,
- 'ytick.labelsize': 8.5,
- 'legend.fontsize': 8,
- 'axes.linewidth': 0.8,
- 'svg.fonttype': 'none',
- 'pdf.fonttype': 42,
- 'ps.fonttype': 42,
- })
- def to_bool(x):
- return str(x).strip().lower() in {'true', '1', 'yes'}
- def load_data(path):
- """Load and prepare data for Venn diagram."""
- print(f"Loading data from: {path}")
- df = pd.read_csv(path)
- print(f"Loaded {len(df)} genes")
- # Create boolean columns for significance
- df['sig_37C'] = df['Sig_37C'].apply(to_bool)
- df['sig_Pulser'] = df['Sig_Pulser'].apply(to_bool)
- df['sig_both'] = df['sig_37C'] & df['sig_Pulser']
- # Get display names
- df['display_name'] = df['Standard_Name'].fillna(df['Gene'])
- return df
- def create_venn(df, output_dir, show_plot=True):
- """Create Venn diagram showing overlap of significant genes."""
- # Get gene lists for each category
- only37 = df[df['sig_37C'] & ~df['sig_Pulser']]['display_name'].tolist()
- both = df[df['sig_both']]['display_name'].tolist()
- onlyp = df[df['sig_Pulser'] & ~df['sig_37C']]['display_name'].tolist()
- # Print statistics
- print(f"\nSignificant genes:")
- print(f" 37°C Heat shock only: {len(only37)} genes")
- print(f" Thermo pulsing only: {len(onlyp)} genes")
- print(f" Both conditions: {len(both)} genes")
- print(f" Total unique significant genes: {len(only37) + len(both) + len(onlyp)}")
- # Create figure
- fig, ax = plt.subplots(figsize=(6.2, 5))
- ax.set_aspect('equal')
- ax.axis('off')
- # Draw circles
- ax.add_patch(Circle((.43, .62), .28,
- facecolor=COLORS['37C'],
- edgecolor=COLORS['37C'],
- alpha=.35, lw=1.5))
- ax.add_patch(Circle((.65, .62), .28,
- facecolor=COLORS['Pulser'],
- edgecolor=COLORS['Pulser'],
- alpha=.35, lw=1.5))
- # Add numbers
- ax.text(.27, .64, len(only37), ha='center', va='center',
- fontsize=18, fontweight='bold')
- ax.text(.54, .64, len(both), ha='center', va='center',
- fontsize=18, fontweight='bold')
- ax.text(.81, .64, len(onlyp), ha='center', va='center',
- fontsize=18, fontweight='bold')
- # Add labels
- ax.text(.23, .94, 'Heat shock (37 °C)', ha='center', fontweight='bold')
- ax.text(.84, .94, 'Thermal pulsing', ha='center', fontweight='bold')
- # Add gene lists (truncated if too long)
- max_genes_display = 10
- gene_lists = []
- for gene_list in [only37, both, onlyp]:
- if len(gene_list) > max_genes_display:
- gene_lists.append(' · '.join(gene_list[:max_genes_display]) + f'\n... and {len(gene_list) - max_genes_display} more')
- elif len(gene_list) > 0:
- gene_lists.append(' · '.join(gene_list))
- else:
- gene_lists.append('(none)')
- ax.text(.18, .22, '37 °C only\n' + gene_lists[0],
- ha='center', va='top', fontsize=8)
- ax.text(.54, .22, 'Both\n' + gene_lists[1],
- ha='center', va='top', fontsize=8)
- ax.text(.88, .22, 'Pulser only\n' + gene_lists[2],
- ha='center', va='top', fontsize=8)
- # Add statistics
- union = len(only37) + len(both) + len(onlyp)
- jaccard = len(both) / union if union > 0 else 0
- ax.text(.5, .04, f'Union = {union}; Jaccard index = {jaccard:.2f}',
- ha='center', fontsize=9)
- ax.set_xlim(0, 1.05)
- ax.set_ylim(0, 1)
- ax.set_title('Overlap of FDR-significant genes', fontweight='bold')
- plt.tight_layout()
- # Show in notebook
- if show_plot:
- plt.show()
- # Save figures
- output_stem = Path(output_dir) / '03_overlap_venn'
- for ext in ['svg', 'png']:
- output_path = output_stem.with_suffix(f'.{ext}')
- plt.savefig(output_path, dpi=300, bbox_inches='tight')
- print(f"✅ Saved: {output_path}")
- plt.close()
- return fig
- # ── Main execution ──────────────────────────────────────────────────────
- # Setup
- configure()
- # Load data
- df = load_data(INPUT_FILE)
- # Create Venn diagram
- fig = create_venn(df, OUTPUT_DIR, show_plot=True)
- print(f"\n✅ All done! Venn diagram saved in: {OUTPUT_DIR}")
- # %% [markdown]
- # ## 5f
- # %%
- # This magic command makes plots appear in the notebook
- %matplotlib inline
- """
- 01_boxplot_verdict_vs_effectsize
- --------------------------------
- Broken-axis violin + box plot: EvoAge verdict vs max signed thermal-protection score.
- INPUT : /storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/01_boxplot_verdict_vs_effectsize/input_boxplot_verdict_data.csv
- OUTPUT: boxplot_verdict_vs_effectsize_brokenaxis.svg / .png / .pdf
- Y-axis: ms = score from the condition (37C or Pulser vs 30C) with larger |value|,
- sign kept. Positive = protective; negative = sensitising.
- Axis break at 1.6 isolates CAX4 (+2.55) and DTD1 (+2.47) from the main distribution.
- """
- import warnings
- warnings.filterwarnings('ignore')
- import pandas as pd
- import numpy as np
- import os
- import matplotlib.pyplot as plt
- from matplotlib.lines import Line2D
- # ── Configuration ──────────────────────────────────────────────────────
- # Input file path
- INPUT_FILE = '/storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/01_boxplot_verdict_vs_effectsize/input_boxplot_verdict_data.csv'
- # Output directory (same as input file location)
- OUTPUT_DIR = os.path.dirname(INPUT_FILE)
- DPI = 600
- # ── Load data ──────────────────────────────────────────────────────────
- print(f"Loading data from: {INPUT_FILE}")
- df = pd.read_csv(INPUT_FILE)
- print(f"Loaded {len(df)} genes")
- # ── Prepare data ──────────────────────────────────────────────────────
- groups = ['support', 'partial_support', 'weak_support', 'no_support']
- data = [df[df['vg'] == g]['ms'].dropna().values for g in groups]
- n = [len(d) for d in data]
- labels = ['Support', 'Partial\nsupport', 'Weak\nsupport', 'No\nsupport']
- colors = ['#1b5e20', '#66bb6a', '#fff9c4', '#d32f2f']
- # Significant genes mapping
- sig_genes = {'YGR036C': 'CAX4', 'YDL219W': 'DTD1', 'YDR098C': 'GRX3',
- 'YGR155W': 'CYS4', 'YDR226W': 'ADK1', 'YBR035C': 'PDX3',
- 'YBL099W': 'ATP1', 'YGL012W': 'ERG4', 'YOR316C': 'COT1',
- 'YGR087C': 'PDC6'}
- # Get positions for significant genes
- key_pos = {}
- for gene, nm in sig_genes.items():
- r = df[df['Gene'] == gene]
- if len(r) > 0:
- r = r.iloc[0]
- key_pos[nm] = (groups.index(r['vg']) if r['vg'] in groups else 0, r['ms'])
- BREAK = 1.6
- # ── Create figure ──────────────────────────────────────────────────────
- fig = plt.figure(figsize=(9.2, 8.4))
- gs = fig.add_gridspec(2, 1, height_ratios=[1.1, 4.2], hspace=0.06)
- # ── TOP PANEL: extreme outliers (> BREAK) ────────────────────────────────
- ax_top = fig.add_subplot(gs[0])
- for i, (d, color) in enumerate(zip(data, colors)):
- hi = d[d > BREAK]
- if len(hi):
- ax_top.scatter([i] * len(hi), hi, s=60, c=color, edgecolors='black',
- linewidth=1.2, zorder=5)
- # Label CAX4 and DTD1 in top panel
- ax_top.annotate('CAX4', key_pos['CAX4'],
- xytext=(key_pos['CAX4'][0] - 0.25, key_pos['CAX4'][1]),
- fontsize=10, fontweight='bold', ha='right', va='center',
- bbox=dict(boxstyle='round,pad=0.2', facecolor='white',
- edgecolor='#999', linewidth=0.6, alpha=0.95), zorder=10)
- ax_top.annotate('DTD1', key_pos['DTD1'],
- xytext=(key_pos['DTD1'][0] + 0.25, key_pos['DTD1'][1] - 0.08),
- fontsize=10, fontweight='bold', ha='left', va='center',
- bbox=dict(boxstyle='round,pad=0.2', facecolor='white',
- edgecolor='#999', linewidth=0.6, alpha=0.95), zorder=10)
- ax_top.set_ylim(BREAK - 0.05, 2.8)
- ax_top.set_xlim(-1.2, 4.5)
- ax_top.set_xticks([])
- ax_top.spines['top'].set_visible(False)
- ax_top.spines['right'].set_visible(False)
- ax_top.spines['bottom'].set_visible(False)
- ax_top.tick_params(labelsize=9)
- ax_top.yaxis.set_major_locator(plt.MultipleLocator(0.5))
- # ── BOTTOM PANEL: main distribution (<= BREAK) ───────────────────────────
- ax = fig.add_subplot(gs[1])
- # Violin plots
- for i, (d, color) in enumerate(zip(data, colors)):
- d_below = d[d <= BREAK]
- if len(d_below) > 0:
- vp = ax.violinplot(d_below, positions=[i], showextrema=False, widths=0.75)
- for body in vp['bodies']:
- body.set_facecolor(color)
- body.set_alpha(0.28)
- body.set_edgecolor('none')
- # Box plots
- bp = ax.boxplot([d[d <= BREAK] for d in data], positions=[0, 1, 2, 3], widths=0.28,
- patch_artist=True, medianprops=dict(color='black', linewidth=2.2),
- flierprops=dict(marker='o', markerfacecolor='#616161', markersize=4,
- markeredgecolor='#424242', markeredgewidth=0.6, alpha=0.8),
- whiskerprops=dict(linewidth=1.4), capprops=dict(linewidth=1.4))
- for patch, color in zip(bp['boxes'], colors):
- patch.set_facecolor(color)
- patch.set_alpha(0.75)
- patch.set_edgecolor('black')
- patch.set_linewidth(1.6)
- ax.axhline(0, color='black', linewidth=0.6, linestyle='-', alpha=0.35)
- # Staggered labels for all 8 bottom-panel significant genes
- bottom_labels = [
- ('GRX3', 0, 1.481, -0.55, 0.02), ('CYS4', 0, 1.247, 0.55, 0.02),
- ('ADK1', 0, 1.159, -0.55, -0.02), ('PDX3', 0, 1.034, 0.55, -0.02),
- ('ATP1', 2, 0.808, -0.55, 0.02), ('ERG4', 2, 0.510, 0.55, 0.02),
- ('PDC6', 2, 0.459, -0.55, -0.02), ('COT1', 2, -0.510, 0.55, -0.02),
- ]
- for nm, gi, val, dx, dy in bottom_labels:
- ax.annotate(nm, xy=(gi, val), xytext=(gi + dx, val + dy),
- fontsize=9, fontweight='bold',
- ha='right' if dx < 0 else 'left', va='center',
- arrowprops=dict(arrowstyle='-', color='#333', lw=0.7,
- shrinkA=2, shrinkB=2, alpha=0.6),
- bbox=dict(boxstyle='round,pad=0.2', facecolor='white',
- edgecolor='#999', linewidth=0.5, alpha=0.9), zorder=10)
- ax.set_ylim(-0.75, BREAK + 0.1)
- ax.set_xlim(-1.2, 4.5)
- ax.set_xticks([0, 1, 2, 3])
- ax.set_xticklabels([f'{l}\n(n={n[i]})' for i, l in enumerate(labels)],
- fontsize=11, fontweight='bold')
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- ax.tick_params(labelsize=9)
- ax.yaxis.set_major_locator(plt.MultipleLocator(0.25))
- # ── BREAK SYMBOLS ────────────────────────────────────────────────────────
- for ax_i in [ax_top, ax]:
- ax_i.plot([-0.03, 0.03], [-0.02, 0.02], transform=ax_i.transAxes,
- color='black', linewidth=1.2, clip_on=False)
- ax_i.plot([-0.03, 0.03], [-0.10, -0.06], transform=ax_i.transAxes,
- color='black', linewidth=1.2, clip_on=False)
- ax.set_ylabel('Max signed thermal-protection score', fontsize=11, fontweight='bold')
- fig.suptitle('EvoAge verdict vs wet-lab effect size — all significant genes labelled',
- fontsize=13, fontweight='bold', y=0.995)
- fig.text(0.5, 0.012,
- 'Y-axis: thermal-protection score (KO - control difference) from the condition '
- '(37 °C vs 30 °C OR Pulser vs 30 °C) with larger absolute value, sign kept.\n'
- 'Positive = knockout improves survival (protective); negative = knockout reduces '
- 'survival (sensitive). Axis break at 1.6. All 10 wet-lab significant genes labelled.',
- ha='center', fontsize=8.5, style='italic', color='#444')
- # Legend
- leg = [Line2D([0], [0], marker='s', color='w', markerfacecolor=c,
- markeredgecolor='black' if c not in ['#fff9c4', '#66bb6a'] else '#555',
- markersize=12, label=l)
- for l, c in [('Support', '#1b5e20'), ('Partial support', '#66bb6a'),
- ('Weak support', '#fff9c4'), ('No support', '#d32f2f')]]
- ax.legend(handles=leg, fontsize=9, loc='upper right', title='EvoAge verdict',
- title_fontsize=10, bbox_to_anchor=(1.0, 1.06))
- plt.tight_layout()
- # ── Show in notebook ──────────────────────────────────────────────────
- plt.show()
- # ── Save figures ──────────────────────────────────────────────────────
- for ext in ['svg', 'png', 'pdf']:
- output_path = os.path.join(OUTPUT_DIR, f'boxplot_verdict_vs_effectsize_brokenaxis.{ext}')
- plt.savefig(output_path, dpi=DPI, bbox_inches='tight')
- print(f"✅ Saved: {output_path}")
- plt.close()
- # ── Print statistics ──────────────────────────────────────────────────
- print("\n📊 Summary statistics:")
- for i, (group, d) in enumerate(zip(labels, data)):
- d_below = d[d <= BREAK]
- d_above = d[d > BREAK]
- print(f"\n{group.strip()}:")
- print(f" Total: {len(d)} genes")
- print(f" Below break (≤{BREAK}): {len(d_below)} genes")
- print(f" Above break (>{BREAK}): {len(d_above)} genes")
- if len(d) > 0:
- print(f" Median: {np.median(d):.3f}")
- print(f" Mean: {np.mean(d):.3f}")
- print(f"\n✅ All done! Plots saved in: {OUTPUT_DIR}")
- # %% [markdown]
- # ## 5g
- # %%
- # This magic command makes plots appear in the notebook
- %matplotlib inline
- """
- 04_bubble_table
- ---------------
- Integrated evidence matrix (bubble table) for the ten FDR-significant genes:
- - Effect sizes (circle colour, diverging blue-red, -0.5 to +2.5)
- - Statistical evidence (circle size, -log10 q, capped at 40 for display)
- - EvoAge verdict pills
- - Known / Novel / Reject / Dissent evidence counts
- INPUT : /storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/04_bubble_table/input_key_genes_evidence.csv
- OUTPUT: EvoAge_Significant_Gene_Bubble_Table.svg / .png / .pdf in the same directory
- """
- import warnings
- warnings.filterwarnings('ignore')
- import pandas as pd
- import numpy as np
- import os
- import matplotlib.pyplot as plt
- from matplotlib.patches import Circle, Rectangle, FancyBboxPatch
- from matplotlib.lines import Line2D
- from matplotlib.colors import LinearSegmentedColormap
- # ── Configuration ──────────────────────────────────────────────────────
- # Input file path
- INPUT_FILE = '/storage/Arushi/090526_EvoAge/kg_formation/all_figures/yeast-exp/EvoAge_plots/04_bubble_table/input_key_genes_evidence.csv'
- # Output directory (same as input file location)
- OUTPUT_DIR = os.path.dirname(INPUT_FILE)
- DPI = 600
- # ── Load data ──────────────────────────────────────────────────────────
- print(f"Loading data from: {INPUT_FILE}")
- df = pd.read_csv(INPUT_FILE)
- print(f"Loaded {len(df)} key genes")
- # Check required columns
- required_cols = ['Name', 'Score_37C', 'Score_Pulser', 'neglog10q_37C', 'neglog10q_Pulser',
- 'verdict', 'known', 'novel', 'reject', 'dissent']
- missing_cols = [col for col in required_cols if col not in df.columns]
- if missing_cols:
- print(f"Warning: Missing columns: {missing_cols}")
- print("Available columns:", df.columns.tolist())
- # ── Prepare data ──────────────────────────────────────────────────────
- # Order: both / 37C / pulser groups
- order = ['CAX4', 'GRX3', 'ADK1', 'CYS4', 'DTD1', 'ATP1', 'COT1', 'ERG4', 'PDC6', 'PDX3']
- group_of = {'CAX4': 'Both', 'GRX3': 'Both', 'ADK1': 'Both', 'CYS4': 'Both',
- 'DTD1': '37 °C', 'ATP1': '37 °C', 'COT1': '37 °C', 'ERG4': '37 °C',
- 'PDC6': '37 °C', 'PDX3': 'Pulser'}
- # Ensure all genes are in the dataframe
- for gene in order:
- if gene not in df['Name'].values:
- print(f"Warning: Gene {gene} not found in data")
- df['order'] = df['Name'].map({n: i for i, n in enumerate(order)})
- df = df.sort_values('order')
- verdict_colors = {'support': '#5E35B1', 'weak_support': '#00897B', # dark purple, teal
- 'partial_support': '#C0CA33', 'no_support': '#BDBDBD'} # muted yellow, grey
- # ── Create figure ──────────────────────────────────────────────────────
- fig, ax = plt.subplots(figsize=(13, 6.5))
- ax.set_xlim(0, 13)
- ax.set_ylim(-1, len(df) + 1.5)
- # Column centres
- col_score37, col_scoreP = 2.2, 3.6
- col_q37, col_qP = 5.2, 6.6
- col_verdict = 8.1
- col_known, col_novel, col_reject, col_dissent = 9.4, 10.2, 11.0, 11.8
- # Headers
- for x, label in [(col_score37, '37 °C'), (col_scoreP, 'Pulser'),
- (col_q37, '37 °C'), (col_qP, 'Pulser'),
- (col_verdict, 'EvoAge verdict'),
- (col_known, 'Known'), (col_novel, 'Novel'),
- (col_reject, 'Reject'), (col_dissent, 'Dissent')]:
- ax.text(x, len(df) + 0.7, label, ha='center', fontsize=9, fontweight='bold')
- ax.text(2.9, len(df) + 1.3, 'Score', ha='center', fontsize=9, fontweight='bold')
- ax.text(5.9, len(df) + 1.3, 'Statistical evidence', ha='center', fontsize=9, fontweight='bold')
- # Group labels
- group_y = {'Both': 4, '37 °C': 1.5, 'Pulser': -0.2}
- for grp, y in group_y.items():
- ax.text(-0.2, y, grp, ha='right', va='center', fontsize=10, fontweight='bold')
- # Diverging colormap for effect size
- cmap = LinearSegmentedColormap.from_list('div', ['#2166AC', '#F7F7F7', '#B2182B'])
- norm_eff = plt.Normalize(-0.5, 2.5)
- # ── Plot data ──────────────────────────────────────────────────────────
- for i, (_, r) in enumerate(df.iterrows()):
- y = len(df) - 1 - i # top-down
- nm = r['Name']
- ax.text(0.4, y, nm, ha='right', va='center', fontsize=9, fontweight='bold')
- # Effect scores
- for x, score in [(col_score37, r['Score_37C']), (col_scoreP, r['Score_Pulser'])]:
- if pd.notna(score):
- c = cmap(norm_eff(score))
- ax.add_patch(Circle((x, y), 0.28, facecolor=c, edgecolor='black', lw=0.8, zorder=3))
- ax.text(x, y, f'{score:+.2f}', ha='center', va='center', fontsize=7, fontweight='bold')
- # Statistical evidence (-log10 q, capped at 40 for display)
- for x, q in [(col_q37, r['neglog10q_37C']), (col_qP, r['neglog10q_Pulser'])]:
- if pd.notna(q):
- size = min(q, 40)
- radius = 0.1 + 0.25 * (size / 40)
- ax.add_patch(Circle((x, y), radius, facecolor='#7E57C2', alpha=0.5,
- edgecolor='#4527A0', lw=0.8, zorder=3))
- ax.text(x, y, f'{q:.1f}', ha='center', va='center', fontsize=7, fontweight='bold')
- # Verdict pill
- v = r['verdict']
- vc = verdict_colors.get(v, '#BDBDBD')
- ax.add_patch(FancyBboxPatch((col_verdict - 0.55, y - 0.22), 1.1, 0.44,
- boxstyle='round,pad=0.02', facecolor=vc,
- edgecolor='black', lw=0.6, zorder=3))
- text_color = 'white' if v == 'support' else 'black'
- ax.text(col_verdict, y, v.replace('_', ' '), ha='center', va='center',
- fontsize=6.5, fontweight='bold', color=text_color)
- # Evidence counts
- for x, val, color in [(col_known, r['known'], '#C8E6C9'),
- (col_novel, r['novel'], '#E1BEE7'),
- (col_reject, r['reject'], '#B3E5FC'),
- (col_dissent, r['dissent'], '#F8BBD0')]:
- ax.add_patch(Rectangle((x - 0.32, y - 0.26), 0.64, 0.52, facecolor=color,
- edgecolor='black', lw=0.5, zorder=3))
- txt = str(int(val)) if pd.notna(val) else 'NA'
- ax.text(x, y, txt, ha='center', va='center', fontsize=8, fontweight='bold')
- ax.axis('off')
- ax.set_title('Thermal-protection score | Triple support count',
- fontsize=11, fontweight='bold', pad=10)
- # Legend for q-size
- for q in [5, 20, 40]:
- ax.add_patch(Circle((11.0, -0.9), 0.1 + 0.25 * (q / 40), facecolor='#7E57C2',
- alpha=0.5, edgecolor='#4527A0', lw=0.6))
- ax.text(12.1, -0.9, '−log10(BH-FDR q), 40-max cap', fontsize=7, va='center')
- # Add effect size colorbar legend
- cbar_ax = fig.add_axes([0.92, 0.15, 0.02, 0.3])
- cbar = plt.colorbar(plt.cm.ScalarMappable(norm=norm_eff, cmap=cmap),
- cax=cbar_ax, orientation='vertical')
- cbar.set_label('Effect size', fontsize=8)
- cbar.ax.tick_params(labelsize=7)
- plt.tight_layout()
- # ── Show in notebook ──────────────────────────────────────────────────
- plt.show()
- # ── Save figures ──────────────────────────────────────────────────────
- output_stem = os.path.join(OUTPUT_DIR, 'EvoAge_Significant_Gene_Bubble_Table')
- for ext in ['svg', 'png', 'pdf']:
- output_path = f"{output_stem}.{ext}"
- plt.savefig(output_path, dpi=DPI, bbox_inches='tight')
- print(f"✅ Saved: {output_path}")
- plt.close()
- # ── Print summary ──────────────────────────────────────────────────────
- print("\n📊 Bubble Table Summary:")
- print(f"Total key genes: {len(df)}")
- print("\nGenes by group:")
- for grp in ['Both', '37 °C', 'Pulser']:
- count = sum([1 for gene in df['Name'] if group_of.get(gene) == grp])
- print(f" {grp}: {count} genes")
- print("\nVerdict distribution:")
- for verdict in ['support', 'weak_support', 'partial_support', 'no_support']:
- count = sum(df['verdict'] == verdict)
- if count > 0:
- print(f" {verdict}: {count} genes")
- print(f"\n✅ All done! Plots saved in: {OUTPUT_DIR}")
- # %% [markdown]
- # ## Supplementary Figure 1
- # %% [markdown]
- # ## Supp 1a
- # %%
- # Make sure inline plotting is enabled
- %matplotlib inline
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import matplotlib.patches as mpatches
- # Optional: If you don't need circlify anymore, remove the import
- # import circlify
- # -----------------------------------------------------------------------------
- # 1. LOAD DATA
- # -----------------------------------------------------------------------------
- df = pd.read_csv("/storage/Arushi/090526_EvoAge/kg_formation/final_kg_building_3/STATS/NodeType_AllKGs_1to1_and_121_12M.csv")
- # -----------------------------------------------------------------------------
- # 2. PRETTY SHORT LABELS + COLOR PALETTE
- # -----------------------------------------------------------------------------
- short_names = {
- "PMID": "PMID",
- "Mutation": "mutation",
- "ChemicalEntity": "chemical",
- "Protein": "protein",
- "Phenotype": "phenotype",
- "Gene": "gene",
- "Disease": "disease",
- "BiologicalProcess": "biological",
- "AnatomicalEntity": "anatomy",
- "MolecularFunction": "molecular",
- "PlantSpecies": "plant",
- "Tissue": "tissue",
- "Pathway": "pathway",
- "CellularComponent": "cellular",
- "Mirna": "mirna",
- "Species": "species",
- }
- nodetype_colors = {
- "PMID": "#c8d89a",
- "Mutation": "#b0b0b0",
- "ChemicalEntity": "#c9a86a",
- "Protein": "#dcb0c4",
- "Phenotype": "#e8a4b5",
- "Gene": "#6e85b0",
- "Disease": "#a89bc8",
- "BiologicalProcess": "#8ec4de",
- "AnatomicalEntity": "#e8a586",
- "MolecularFunction": "#c9a86a",
- "PlantSpecies": "#a67aa8",
- "Tissue": "#f0a878",
- "Pathway": "#b39ddb",
- "CellularComponent": "#f5d896",
- "Mirna": "#b06ba0",
- "Species": "#dcdc82",
- }
- # -----------------------------------------------------------------------------
- # 3. CREATE THE PLOT
- # -----------------------------------------------------------------------------
- # Clear any existing figures
- plt.close('all')
- # Create figure and axis
- fig, ax1 = plt.subplots(1, 1, figsize=(12, 8))
- # =============================================================================
- # BAR PLOT — EvoAge_121_12M
- # =============================================================================
- evo = df[["NodeType", "EvoAge_121_12M"]].copy()
- evo = evo[evo["EvoAge_121_12M"] > 0].sort_values("EvoAge_121_12M", ascending=False).reset_index(drop=True)
- evo["log"] = np.log10(evo["EvoAge_121_12M"] + 1)
- x = np.arange(len(evo))
- bar_colors = [nodetype_colors.get(nt, "#cccccc") for nt in evo["NodeType"]]
- # Create the bars
- bars = ax1.bar(x, evo["log"], color=bar_colors, edgecolor="black", linewidth=0.4, width=0.75)
- # Count labels rotated 90° inside/on each bar
- for xi, (_, row) in zip(x, evo.iterrows()):
- ax1.text(xi, row["log"] * 0.5, f"{int(row['EvoAge_121_12M'])}",
- ha="center", va="center", rotation=90, fontsize=9, color="black")
- # Set labels and title
- ax1.set_xticks(x)
- ax1.set_xticklabels([short_names.get(nt, nt) for nt in evo["NodeType"]],
- rotation=90, fontsize=10)
- ax1.set_ylabel(r"$\log_{10}$(count)", fontsize=13)
- ax1.set_title("Node count (EvoAge)", fontsize=15, pad=10)
- # Remove top and right spines
- ax1.spines["top"].set_visible(False)
- ax1.spines["right"].set_visible(False)
- ax1.tick_params(axis="both", which="major", length=5, width=0.8, color="black", labelsize=10)
- ax1.grid(False)
- # Adjust layout
- plt.tight_layout()
- # -----------------------------------------------------------------------------
- # 4. DISPLAY THE PLOT
- # -----------------------------------------------------------------------------
- # Show the plot
- plt.show()
- # -----------------------------------------------------------------------------
- # 5. SAVE AS PDF (after showing)
- # -----------------------------------------------------------------------------
- # Save the figure
- fig.savefig("FIG1/Fig1_supp_evoage_bar.pdf", format="pdf", bbox_inches="tight")
- print("Plot saved successfully!")
- # %% [markdown]
- # ## Supp 1b
- # %%
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- from matplotlib.path import Path
- from matplotlib.patches import PathPatch, Wedge, Rectangle
- import matplotlib as mpl
- # ------------------------------------------------------------------
- # 0. Editable SVG text
- # ------------------------------------------------------------------
- mpl.rcParams['svg.fonttype'] = 'none'
- mpl.rcParams['font.family'] = 'DejaVu Sans'
- # ------------------------------------------------------------------
- # 1. Load & parse
- # ------------------------------------------------------------------
- CSV_PATH = 'RelationType_AllKGs_1to1_and_121_12M.csv'
- VALUE_COL = 'EvoAge_121_12M'
- df = pd.read_csv(CSV_PATH)
- df = df[df['Relation'] != 'Total Triples'].copy()
- df = df[df[VALUE_COL] > 0].reset_index(drop=True)
- def parse_relation(relation):
- parts = relation.split('_')
- if len(parts) == 2:
- return parts[0], None, parts[1], 'entity'
- else:
- # first token = source, LAST token = target, everything in between = verb
- return parts[0], '_'.join(parts[1:-1]), parts[-1], 'verb'
- parsed = df['Relation'].apply(parse_relation)
- df['src'] = parsed.apply(lambda x: x[0])
- df['verb'] = parsed.apply(lambda x: x[1])
- df['tgt'] = parsed.apply(lambda x: x[2])
- df['edge_type'] = parsed.apply(lambda x: x[3])
- # camel-case -> "spaced lower case" for nice legend text, e.g.
- # "NegativelyAssociatedWith" -> "negatively associated with"
- def humanize(camel):
- out = []
- for ch in camel:
- if ch.isupper() and out:
- out.append(' ')
- out.append(ch.lower())
- return ''.join(out)
- # ------------------------------------------------------------------
- # 2. log10 weight used for ribbon/arc geometry (raw count kept for labels)
- # ------------------------------------------------------------------
- df['raw'] = df[VALUE_COL]
- df['logw'] = np.log10(df['raw'] + 1)
- # ------------------------------------------------------------------
- # 3. Entities & node ordering
- # ------------------------------------------------------------------
- entities = sorted(set(df['src']) | set(df['tgt']))
- idx = {e: i for i, e in enumerate(entities)}
- n = len(entities)
- # total log-weighted flux touching each entity (for arc size)
- flux = np.zeros(n)
- for _, row in df.iterrows():
- flux[idx[row['src']]] += row['logw']
- flux[idx[row['tgt']]] += row['logw']
- order = np.argsort(-flux)
- entities = [entities[i] for i in order]
- idx = {e: i for i, e in enumerate(entities)}
- flux = flux[order]
- # ------------------------------------------------------------------
- # 4. Colour palettes
- # - entity palette: pastel, one colour per entity (used for entity-entity
- # ribbons AND node wedges)
- # - relation-type palette: separate, more saturated colours, used ONLY for
- # verb-qualified ribbons (Promotes / Inhibits / etc.)
- # ------------------------------------------------------------------
- # Colours matched (pixel-sampled) from the reference "EvoAge (Node count)"
- # bar chart, so the entity colours are consistent across figures.
- entity_hex = {
- 'PMID': '#C8D89A', # olive-green
- 'Mutation': '#B0B0B0', # grey
- 'ChemicalEntity': '#C9A86A', # tan / gold-brown
- 'Protein': '#DCB0C4', # pink
- 'Phenotype': '#E8A4B5', # rose pink
- 'Gene': '#6E85B0', # dark blue / navy
- 'AnatomicalEntity': '#E8A586', # salmon / orange
- 'Disease': '#A89BC8', # light purple
- 'BiologicalProcess': '#8EC4DE', # sky blue
- 'MolecularFunction': '#B08968', # muted brown (reference re-used the
- # chemical tan here; shifted slightly
- # darker so every entity stays
- # distinguishable, per your note)
- 'PlantSpecies': '#A67AA8', # plum
- 'Tissue': '#F0A878', # orange
- 'Pathway': '#B39DDB', # lavender-purple
- 'CellularComponent': '#F5D896', # cream / pale yellow
- 'Mirna': '#B06BA0', # magenta-plum
- 'Species': '#DCDC82', # yellow-green
- 'Nodes': '#C0504D', # brick red (not in the reference
- # chart, so a new but equally
- # contrasting colour was added)
- }
- node_color = {e: entity_hex.get(e, '#CCCCCC') for e in entities}
- # Relation-type (verb-qualified edge) palette: kept intentionally as
- # saturated/dark "jewel tones" -- a different visual register from the
- # pastel entity colours above -- so verb-qualified ribbons never get
- # confused with a plain entity-entity ribbon, even where hue families
- # are close (e.g. reds vs pinks, greens vs olive).
- verb_palette = {
- 'AssociatedWith': '#073B3A', # deep teal
- 'NoEffect': '#6A00FF', # vivid violet
- 'Promotes': '#003566', # dark navy
- 'Inhibits': '#9D0208', # dark red
- 'NegativelyAssociatedWith': '#2B9348', # deep green
- 'PositivelyAssociatedWith': '#BC6C25', # burnt amber
- 'NotAssociatedWith': '#4A4E69', # slate grey-purple
- }
- def edge_color(row):
- if row['edge_type'] == 'entity':
- return node_color[row['src']]
- return verb_palette.get(row['verb'], '#999999')
- df['color'] = df.apply(edge_color, axis=1)
- # short display names for entity labels (match reference style, e.g.
- # "BiologicalProcess" -> "Biological")
- short_name = {
- 'BiologicalProcess': 'Biological',
- 'CellularComponent': 'Cellular',
- 'MolecularFunction': 'Molecular',
- 'AnatomicalEntity': 'Anatomical',
- 'ChemicalEntity': 'Chemical',
- 'PlantSpecies': 'Plant Species',
- }
- def short(e):
- return short_name.get(e, e)
- # ------------------------------------------------------------------
- # 5. Layout: angles for each node's arc (deg), clockwise from top
- # ------------------------------------------------------------------
- GAP_DEG = 2.4
- total_gap = GAP_DEG * n
- available_deg = 360 - total_gap
- frac = flux / flux.sum()
- node_span = frac * available_deg
- start_ang = 90.0
- node_start, node_end = {}, {}
- cursor = start_ang
- for e, span in zip(entities, node_span):
- node_start[e] = cursor - span
- node_end[e] = cursor
- cursor = cursor - span - GAP_DEG
- node_center = {e: (node_start[e] + node_end[e]) / 2 for e in entities}
- # ------------------------------------------------------------------
- # 6. Slots (one per ribbon end) sized by log-weight
- # ------------------------------------------------------------------
- slots = {e: [] for e in entities}
- for edge_id, row in df.iterrows():
- s, t, w = row['src'], row['tgt'], row['logw']
- slots[s].append(dict(partner=t, weight=w, edge_id=edge_id, end='src'))
- slots[t].append(dict(partner=s, weight=w, edge_id=edge_id, end='tgt'))
- for e in entities:
- def sort_key(sl):
- p = sl['partner']
- return -1000 if p == e else node_center[p]
- slots[e].sort(key=sort_key)
- edge_geom = {}
- for e in entities:
- s_list = slots[e]
- wsum = sum(sl['weight'] for sl in s_list)
- if wsum == 0:
- continue
- span = node_end[e] - node_start[e]
- cur = node_start[e]
- for sl in s_list:
- w = span * (sl['weight'] / wsum)
- a0, a1 = cur, cur + w
- cur = a1
- edge_geom.setdefault(sl['edge_id'], {})[sl['end']] = (a0, a1)
- # ------------------------------------------------------------------
- # 7. Drawing helpers
- # ------------------------------------------------------------------
- R_IN = 1.0
- R_OUT = 1.09
- N_ARC_PTS = 40
- def arc_points(a0, a1, r=R_IN):
- angs = np.linspace(np.radians(a0), np.radians(a1), N_ARC_PTS)
- return np.column_stack([r * np.cos(angs), r * np.sin(angs)])
- def ribbon_path(srcA, srcB, tgtA, tgtB):
- p_src = arc_points(srcA, srcB)
- p_tgt = arc_points(tgtA, tgtB)
- verts = [p_src[0]]
- codes = [Path.MOVETO]
- for p in p_src[1:]:
- verts.append(p); codes.append(Path.LINETO)
- c1, c2 = p_src[-1] * 0.35, p_tgt[0] * 0.35
- verts += [c1, c2, p_tgt[0]]; codes += [Path.CURVE4]*3
- for p in p_tgt[1:]:
- verts.append(p); codes.append(Path.LINETO)
- c3, c4 = p_tgt[-1] * 0.35, p_src[0] * 0.35
- verts += [c3, c4, p_src[0]]; codes += [Path.CURVE4]*3
- return Path(verts, codes)
- # ------------------------------------------------------------------
- # 8. Plot
- # ------------------------------------------------------------------
- fig, ax = plt.subplots(figsize=(14, 12), subplot_kw=dict(aspect='equal'))
- ax.set_xlim(-1.85, 2.35)
- ax.set_ylim(-1.55, 1.65)
- ax.axis('off')
- edges_sorted = df.sort_values('raw', ascending=False).index.tolist()
- for eid in edges_sorted:
- row = df.loc[eid]
- if eid not in edge_geom or 'src' not in edge_geom[eid] or 'tgt' not in edge_geom[eid]:
- continue
- a0, a1 = edge_geom[eid]['src']
- b0, b1 = edge_geom[eid]['tgt']
- path = ribbon_path(a0, a1, b0, b1)
- patch = PathPatch(path, facecolor=row['color'], edgecolor='none', alpha=0.65,
- linewidth=0, zorder=1)
- ax.add_patch(patch)
- # node wedges
- for e in entities:
- w = Wedge((0, 0), R_OUT, node_start[e], node_end[e], width=R_OUT - R_IN,
- facecolor=node_color[e], edgecolor='white', linewidth=1.2, zorder=3)
- ax.add_patch(w)
- # raw-count numbers next to each ribbon slot
- for eid, ends in edge_geom.items():
- raw = df.loc[eid, 'raw']
- label = f'{int(raw):,}'
- for end_key in ('src', 'tgt'):
- if end_key not in ends:
- continue
- a0, a1 = ends[end_key]
- mid = np.radians((a0 + a1) / 2)
- r_txt = R_OUT + 0.025
- x, y = r_txt * np.cos(mid), r_txt * np.sin(mid)
- rot = np.degrees(mid)
- ha = 'left'
- if 90 < (np.degrees(mid) % 360) < 270:
- rot += 180
- ha = 'right'
- ax.text(x, y, label, rotation=rot, rotation_mode='anchor',
- ha=ha, va='center', fontsize=6.3, zorder=5)
- # entity name labels (outer ring)
- for e in entities:
- ang = np.radians(node_center[e])
- r_label = R_OUT + 0.16
- x, y = r_label * np.cos(ang), r_label * np.sin(ang)
- rot = node_center[e]
- ha = 'left'
- if 90 < node_center[e] % 360 < 270:
- rot += 180
- ha = 'right'
- ax.text(x, y, short(e), rotation=rot, rotation_mode='anchor',
- ha=ha, va='center', fontsize=11, fontweight='medium', zorder=4)
- # "log10(count)" radial axis label (reference-style)
- ax.text(-1.42, 0, r'$\log_{10}(\mathrm{count})$', rotation=90, rotation_mode='anchor',
- ha='center', va='center', fontsize=13)
- ax.set_title('EvoAge Knowledge Graph — Entity Relation Chord Diagram\n'
- 'relation counts (EvoAge, 121:12M) | ribbon width \u221d log$_{10}$(count)',
- fontsize=13, pad=10)
- # ------------------------------------------------------------------
- # 9. Legend: entities block + relation-type block
- # ------------------------------------------------------------------
- leg_x = 1.42
- leg_y = 1.55
- box_w, box_h, gap = 0.075, 0.055, 0.018
- ax.text(leg_x, leg_y + 0.06, 'Entities', fontsize=12.5, fontweight='bold', va='center')
- y = leg_y
- for e in entities:
- ax.add_patch(Rectangle((leg_x, y - box_h/2), box_w, box_h,
- facecolor=node_color[e], edgecolor='none'))
- ax.text(leg_x + box_w + 0.03, y, short(e), fontsize=10.5, va='center')
- y -= (box_h + gap)
- y -= 0.04
- ax.text(leg_x, y + 0.06, 'Relation type\n(qualified edges, separate colour scheme)',
- fontsize=11, fontweight='bold', va='center')
- y -= 0.02
- for verb, col in verb_palette.items():
- if verb not in df.loc[df['edge_type']=='verb', 'verb'].unique():
- continue
- ax.add_patch(Rectangle((leg_x, y - box_h/2), box_w, box_h,
- facecolor=col, edgecolor='none'))
- ax.text(leg_x + box_w + 0.03, y, humanize(verb), fontsize=10.5, va='center')
- y -= (box_h + gap)
- plt.tight_layout()
- plt.savefig('EvoAge_ChordDiagram_log10_matched.svg', format='svg', bbox_inches='tight')
- plt.savefig('EvoAge_ChordDiagram_log10_matched_preview.png', format='png', dpi=170, bbox_inches='tight')
- print("Saved EvoAge_ChordDiagram_log10_matched.svg and preview png")
- # %% [markdown]
- # ## Supp 1c
- # %%
- import pandas as pd
- import plotly.graph_objects as go
- # Load data
- df = pd.read_csv('orthology-mapping.csv')
- print("Data loaded successfully!")
- print(df)
- # Prepare data
- species_list = df['Species'].tolist()
- labels = species_list + ['1-to-1\northolog', '1-to-M\northologs', 'Unmapped']
- source = []
- target = []
- value = []
- species_indices = {species: idx for idx, species in enumerate(species_list)}
- category_indices = {
- 'OneToOne': len(species_list),
- 'OneToMany': len(species_list) + 1,
- 'Unmapped': len(species_list) + 2
- }
- for idx, row in df.iterrows():
- species = row['Species']
- species_idx = species_indices[species]
- source.append(species_idx)
- target.append(category_indices['OneToOne'])
- value.append(row['OneToOne'])
- source.append(species_idx)
- target.append(category_indices['OneToMany'])
- value.append(row['OneToMany'])
- source.append(species_idx)
- target.append(category_indices['Unmapped'])
- value.append(row['Unmapped'])
- total_1to1 = df['OneToOne'].sum()
- total_1toM = df['OneToMany'].sum()
- total_unmapped = df['Unmapped'].sum()
- # Cleaner version with gray links
- fig = go.Figure(data=[go.Sankey(
- node=dict(
- pad=30,
- thickness=25,
- line=dict(color="white", width=1),
- label=labels,
- color=['#3498DB', '#E67E22', '#2ECC71', '#E74C3C', '#9B59B6',
- '#2980B9', '#F39C12', '#C0392B'],
- x=[0.05, 0.05, 0.05, 0.05, 0.05, 0.82, 0.82, 0.82],
- y=[0.15, 0.32, 0.49, 0.66, 0.83, 0.25, 0.55, 0.80]
- ),
- link=dict(
- source=source,
- target=target,
- value=value,
- color='rgba(180, 180, 180, 0.5)' # Gray semi-transparent links
- )
- )])
- fig.update_layout(
- title=dict(
- text="<b>Ortholog Gene Mapping from Other Species to Human</b>",
- font=dict(size=22, family="Arial", color="#2C3E50"),
- x=0.5
- ),
- width=1200,
- height=700,
- margin=dict(t=80, b=120, l=60, r=60),
- plot_bgcolor='white',
- paper_bgcolor='white'
- )
- # Add annotations
- fig.add_annotation(
- x=0.5,
- y=-0.10,
- xref="paper",
- yref="paper",
- text=(f"<b>Summary</b> | "
- f"<span style='color:#2980B9;'>●</span> 1-to-1: <b>{total_1to1:,}</b> | "
- f"<span style='color:#F39C12;'>●</span> 1-to-M: <b>{total_1toM:,}</b> | "
- f"<span style='color:#C0392B;'>●</span> Unmapped: <b>{total_unmapped:,}</b>"),
- showarrow=False,
- font=dict(size=14),
- align="center"
- )
- fig.show()
- fig.write_html("sankey_ortholog_mapping_simple.html")
- print("\n✅ Simple Sankey diagram saved as 'sankey_ortholog_mapping_simple.html'")
- # %% [markdown]
- # ## Supp 1e
- # %%
- import pandas as pd
- import matplotlib.pyplot as plt
- import numpy as np
- from matplotlib.patches import Rectangle
- import matplotlib.colors as mcolors
- # Read the CSV file
- df = pd.read_csv('ALL_data_Source_Triple_source_databases_KG.csv')
- # Filter for aging data sources only (KG_Type == 'Aging')
- aging_df = df[df['KG_Type'] == 'Aging'].copy()
- # Define all columns to check (all relationship types)
- all_columns = aging_df.columns.tolist()
- # Remove non-data columns
- exclude_cols = ['Source', 'KG_Type', 'Total', 'Species_AssociatedWith', 'PlantSpecies_ChemicalEntity',
- 'PlantSpecies_Disease', 'PMID_CellularComponent', 'PMID_ChemicalEntity', 'PMID_Disease',
- 'PMID_Protein', 'PMID_Tissue', 'Phenotype_BiologicalProcess', 'Phenotype_CellularComponent',
- 'Phenotype_MolecularFunction', 'Phenotype_Protein', 'Phenotype_Phenotype', 'Phenotype_Gene',
- 'Phenotype_Disease', 'Phenotype_ChemicalEntity', 'Mirna_Gene']
- columns_to_plot = [col for col in all_columns if col not in exclude_cols and col != 'Total']
- # Filter to columns that have some data in aging sources
- valid_columns = []
- for col in columns_to_plot:
- if col in aging_df.columns and aging_df[col].sum() > 0:
- valid_columns.append(col)
- # Prepare the data matrix
- aging_data = aging_df[['Source'] + valid_columns].copy()
- aging_data.set_index('Source', inplace=True)
- # Remove sources that have no data in these columns
- aging_data = aging_data[aging_data.sum(axis=1) > 0]
- # Replace 0 with NaN (we don't want to display zeros)
- aging_data = aging_data.replace(0, np.nan)
- # Apply log10 transformation (adding small constant to avoid log(0) issues)
- # We'll use log10(count + 1) for the color mapping
- log_data = np.log10(aging_data.values + 1)
- # The +1 ensures log(1)=0 for zero values, but we'll mask NaNs anyway
- # Create the plot with the specific color gradient
- fig, ax = plt.subplots(figsize=(25, 8))
- # Define the colormap matching the screenshot (blue to red)
- # This creates a diverging colormap from blue through white to red
- cmap = plt.cm.viridis
- # Create the heatmap
- # We use log_data for color mapping but keep original values for text
- im = ax.imshow(log_data, cmap=cmap, aspect='auto',
- vmin=np.nanmin(log_data), vmax=np.nanmax(log_data))
- # Set ticks and labels
- ax.set_xticks(np.arange(len(aging_data.columns)))
- ax.set_yticks(np.arange(len(aging_data.index)))
- ax.set_xticklabels(aging_data.columns, rotation=90, ha='center', fontsize=9)
- ax.set_yticklabels(aging_data.index, fontsize=10)
- # Add colored boxes with borders for each cell
- for i in range(len(aging_data.index)):
- for j in range(len(aging_data.columns)):
- value = aging_data.iloc[i, j]
- if not pd.isna(value):
- # Add rectangle border
- rect = Rectangle((j-0.5, i-0.5), 1, 1, fill=False, edgecolor='black', linewidth=1.5)
- ax.add_patch(rect)
- # Format the value
- if value >= 1000000:
- text = f'{value/1000000:.1f}M'
- elif value >= 1000:
- text = f'{value/1000:.1f}K'
- else:
- text = f'{int(value)}'
- # Choose text color based on background
- log_val = log_data[i, j]
- if log_val > np.nanmax(log_data) * 0.4:
- color = 'white'
- else:
- color = 'black'
- ax.text(j, i, text, ha='center', va='center', color=color, fontsize=7, weight='bold')
- # Add colorbar with log scale labels
- cbar = plt.colorbar(im, ax=ax, shrink=0.8)
- cbar.set_label('log₁₀(count + 1)', fontsize=12)
- # Add legend for log values
- log_ticks = np.arange(0, np.nanmax(log_data) + 0.5, 0.5)
- cbar.set_ticks(log_ticks)
- cbar.set_ticklabels([f'{tick:.1f}' for tick in log_ticks])
- # Set title
- ax.set_title('Aging Data Sources - Relationship Types (Log Scale)', fontsize=16, pad=20)
- # Adjust layout
- plt.tight_layout()
- # Save the figure (optional)
- # plt.savefig('aging_data_heatmap.png', dpi=300, bbox_inches='tight')
- plt.show()
- # Print summary statistics
- print("Aging Data Sources Summary:")
- print("=" * 60)
- print(f"Number of aging data sources: {len(aging_data.index)}")
- print(f"Data sources: {', '.join(aging_data.index)}")
- print(f"\nNumber of relationship types with data: {len(aging_data.columns)}")
- print("\nTop relationship types by total count:")
- totals = aging_data.sum(axis=0).sort_values(ascending=False)
- for col, val in totals.head(10).items():
- print(f" {col}: {int(val):,}")
- # Alternative: Create a version that only shows columns where any source has data
- def create_filtered_heatmap():
- """Create a more focused heatmap with only active relationship types"""
- # Get columns where at least one source has data
- active_columns = []
- for col in valid_columns:
- if aging_df[col].sum() > 0:
- active_columns.append(col)
- # Prepare data
- filtered_data = aging_df[['Source'] + active_columns].copy()
- filtered_data.set_index('Source', inplace=True)
- filtered_data = filtered_data.replace(0, np.nan)
- filtered_data = filtered_data[filtered_data.sum(axis=1) > 0]
- # Log transform
- log_filtered = np.log10(filtered_data.values + 1)
- # Plot
- fig, ax = plt.subplots(figsize=(20, 8))
- im = ax.imshow(log_filtered, cmap='RdYlBu_r', aspect='auto',
- vmin=0, vmax=np.nanmax(log_filtered))
- ax.set_xticks(np.arange(len(filtered_data.columns)))
- ax.set_yticks(np.arange(len(filtered_data.index)))
- ax.set_xticklabels(filtered_data.columns, rotation=90, ha='center', fontsize=8)
- ax.set_yticklabels(filtered_data.index, fontsize=10)
- # Add boxes and text
- for i in range(len(filtered_data.index)):
- for j in range(len(filtered_data.columns)):
- value = filtered_data.iloc[i, j]
- if not pd.isna(value):
- rect = Rectangle((j-0.5, i-0.5), 1, 1, fill=False, edgecolor='black', linewidth=1.5)
- ax.add_patch(rect)
- if value >= 1000000:
- text = f'{value/1000000:.1f}M'
- elif value >= 1000:
- text = f'{value/1000:.1f}K'
- else:
- text = f'{int(value)}'
- log_val = log_filtered[i, j]
- color = 'white' if log_val > np.nanmax(log_filtered) * 0.4 else 'black'
- ax.text(j, i, text, ha='center', va='center', color=color, fontsize=7, weight='bold')
- cbar = plt.colorbar(im, ax=ax, shrink=0.8)
- cbar.set_label('log₁₀(count + 1)', fontsize=12)
- ax.set_title('Aging Data Sources - Active Relationship Types (Log Scale)', fontsize=14, pad=15)
- plt.tight_layout()
- plt.show()
- # Uncomment to see the filtered version
- # create_filtered_heatmap()
- # Create a table version showing only the non-zero values
- def create_table_view():
- """Create a table view similar to the screenshot"""
- # Get the data
- table_data = aging_df[['Source'] + valid_columns].copy()
- table_data.set_index('Source', inplace=True)
- # Only keep columns with data
- cols_with_data = table_data.columns[table_data.sum(axis=0) > 0]
- table_data = table_data[cols_with_data]
- # Only keep rows with data
- rows_with_data = table_data.index[table_data.sum(axis=1) > 0]
- table_data = table_data.loc[rows_with_data]
- # Display as a styled table
- fig, ax = plt.subplots(figsize=(20, len(table_data) * 0.5 + 1))
- ax.axis('off')
- # Create the table
- table = ax.table(cellText=table_data.values.astype(str).tolist(),
- rowLabels=table_data.index,
- colLabels=table_data.columns,
- cellLoc='center',
- rowLoc='center',
- loc='center')
- table.auto_set_font_size(False)
- table.set_fontsize(8)
- table.scale(1.2, 1.5)
- # Color the cells based on values
- for i in range(len(table_data.index)):
- for j in range(len(table_data.columns)):
- value = table_data.iloc[i, j]
- if value > 0:
- # Get cell
- cell = table[(i+1, j)]
- # Color based on log value
- log_val = np.log10(value + 1)
- max_log = np.log10(table_data.max().max() + 1)
- if log_val / max_log > 0.6:
- cell.set_facecolor('#e74c3c') # Red
- cell.set_text_props(color='white')
- elif log_val / max_log > 0.3:
- cell.set_facecolor('#f39c12') # Orange
- cell.set_text_props(color='white')
- else:
- cell.set_facecolor('#3498db') # Blue
- cell.set_text_props(color='white')
- else:
- cell = table[(i+1, j)]
- cell.set_facecolor('#f0f0f0')
- cell.set_text_props(color='gray')
- plt.title('Aging Data Sources - Relationship Counts', fontsize=14, pad=20)
- plt.tight_layout()
- plt.show()
- # Uncomment to see the table view
- # create_table_view()
- # %% [markdown]
- # ## Supp 1f
- # %%
- import pandas as pd
- import matplotlib.pyplot as plt
- import numpy as np
- from matplotlib.patches import Rectangle
- import matplotlib.colors as mcolors
- # Read the CSV file
- df = pd.read_csv('Triple_source_databases_KG.csv')
- # Filter for generalised data sources only (KG_Type == 'Generalised')
- generalised_df = df[df['KG_Type'] == 'Generalised'].copy()
- # Define the exact order of sources from the screenshot
- source_order = [
- 'BOCK',
- 'BindingDB',
- 'BioGrakn',
- 'BioSNAP',
- 'Biogrid', # Note: BioGRID in the list
- 'CKG',
- 'CROssBAR',
- 'Chembl', # Note: ChEMBL in the list
- 'DRKG',
- 'DTINet',
- 'DrugBank',
- 'Evolf', # Note: EvOf in the list
- 'FICE',
- 'FlyBase',
- 'Harmonizome',
- 'Hetionet',
- 'IMPPAT',
- 'MGI_DO', # Note: MGI in the list
- 'MonarchKG',
- 'Mouse_Net', # Note: MouseNet in the list
- 'PharmKG',
- 'PheKnowLator',
- 'Phytohub',
- 'PrimeKG',
- 'SGD',
- 'SMS',
- 'STITCH',
- 'STRING',
- 'TARKG',
- 'TTD',
- 'Worm_Interactome_Database', # Note: WIDB in the list
- 'YeastNet',
- 'ZFIN',
- 'eSLDB',
- 'iBKH',
- 'miRTarBase', # Note: miRtarBase in the list
- 'other_sources'
- ]
- # Create a mapp
all_figures-raidhani.ipynb at commit 6067db0, under MIT · at the source
Overview
- Indraprastha Institute of Information Technology-Delhi (IIIT-Delhi)
- Indraprastha Institute of Information Technology
- University of Cambridge
- Indraprastha Institute of Information Technology-Delhi, (IIIT-Delhi)
- Department of Computational Biology, Indraprastha Institute of Information Technology-Delhi (IIIT-Delhi), Okhla, Phase III, New Delhi, 110020, India
- Indian Institute of Science
- University of Copenhagen
- Indraprastha Institute of Information Technology Delhi
Abstract
Aging research has been advanced largely through the use of model organisms, where short lifespans and genetic tractability enable the systematic discovery of molecular pathways influencing longevity and age-related decline. However, knowledge about aging remains fragmented across species-specific repositories and domain-focused databases, limiting our ability to identify evolutionarily conserved mechanisms and translate findings to human biology. To address this gap, we developed EvoAge, a unified, multi-species knowledge graph that integrates aging-specific and general biomedical resources into a systems-level framework. EvoAge harmonizes 48 public datasets into a graph comprising 1.04 billion triples across six key species. A human-centric orthology framework reconciles more than 80,000 gene entries, expanding accessible organism-level aging knowledge by up to 1,700-fold compared with existing resources. To operationalize the graph for biological reasoning, we optimized knowledge graph embedding models and deployed a large language model (LLM)-assisted agentic interface that supports natural-language querying, link prediction, and hypothesis testing. In internal benchmarking using recent pre-print aging literature, EvoAge significantly outperformed state-of-the-art LLMs in distinguishing biologically plausible from implausible hypotheses. Importantly, EvoAge recommended a previously unrecognized Alzheimer’s disease (AD) mechanism involving nanoscale redistribution of BACE1 within synaptic compartments. We experimentally validated this EvoAge-supported prediction using patient-derived iPSCs carrying a familial PSEN1 mutation, demonstrating disease-associated remodeling of β-secretase, defined by altered localization, nanoscale clustering, and compartment-specific enrichment. We further confirmed the predicted evolutionary conservation of this BACE1–pathology relationship in additional AD systems, including transgenic mice and postmortem human brain tissue.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 21 matches between paragraphs and lines of code.
the-ahuja-lab/EvoAge
6067db0a3ac902952f48933db8e72dac4fd15901, 25 August 2026Availability: 1 check, the latest on 30 September 2026: the link answers
- 30 September 2026: the link answers
421 files
- Backend/
app/ , Python, 1 line__init__.py - Backend/
app/ , Python, 219 linesauth_routes.py - Backend/
app/ , Python, 152 linescrud/ user.py - Backend/
app/ , Python, 40 linesdemo_routes.py - Backend/
app/ , Python, 1,842 lineshypothesis_routes.py - Backend/
app/ , Python, 118 linesmain.py - Backend/
app/ , Python, 765 linesmodel_routes.py - Backend/
app/ , Python, 67 linesmodels/ user.py - Backend/
app/ , Python, 9 linesmodels/ utils_models.py - Backend/
app/ , Python, 914 linesroutes.py - Backend/
app/ , Python, 193 linesuser_routes.py - Backend/
app/ , Python, 1 lineutils/ __init__.py - Backend/
app/ , Python, 212 linesutils/ constants.py - Backend/
app/ , Python, 87 linesutils/ database.py - Backend/
app/ , Python, 497 linesutils/ email_utils.py - Backend/
app/ , Python, 159 linesutils/ environment.py - Backend/
app/ , Python, 1,107 linesutils/ hypothesis_utils.py - Backend/
app/ , Python, 228 linesutils/ model_manager.py - Backend/
app/ , Python, 28 linesutils/ redis_utils.py - Backend/
app/ , Python, 94 linesutils/ schema.py - Backend/
app/ , Python, 85 linesutils/ security.py - Backend/
app/ , Python, 27 linesutils_routes.py - Backend/
deploy.sh , Shell, 11 lines - Backend/
deploy_direst.sh , Shell, 17 lines - Backend/
dgl-ke/ , Python, 24 linespython/ dglke/ __init__.py - Backend/
dgl-ke/ , Python, 75 linespython/ dglke/ convert.py - Backend/
dgl-ke/ , Python, 834 linespython/ dglke/ dataloader/ KGDataset.py - Backend/
dgl-ke/ , Python, 21 linespython/ dglke/ dataloader/ __init__.py - Backend/
dgl-ke/ , Python, 882 linespython/ dglke/ dataloader/ sampler.py - Backend/
dgl-ke/ , Python, 218 linespython/ dglke/ dist_train.py - Backend/
dgl-ke/ , Python, 229 lines, 1 matchpython/ dglke/ eval.py - Backend/
dgl-ke/ , Python, 139 linespython/ dglke/ infer_emb_sim.py - Backend/
dgl-ke/ , Python, 237 linespython/ dglke/ infer_score.py - Backend/
dgl-ke/ , Python, 223 linespython/ dglke/ kvclient.py - Backend/
dgl-ke/ , Python, 181 linespython/ dglke/ kvserver.py - Backend/
dgl-ke/ , Python, 21 linespython/ dglke/ models/ __init__.py - Backend/
dgl-ke/ , Python, 161 linespython/ dglke/ models/ base_loss.py - Backend/
dgl-ke/ , Python, 680 linespython/ dglke/ models/ general_models.py - Backend/
dgl-ke/ , Python, 446 linespython/ dglke/ models/ infer.py - Backend/
dgl-ke/ , Python, 978 linespython/ dglke/ models/ ke_model.py - Backend/
dgl-ke/ , Python, 18 linespython/ dglke/ models/ mxnet/ __init__.py - Backend/
dgl-ke/ , Python, 45 linespython/ dglke/ models/ mxnet/ loss.py - Backend/
dgl-ke/ , Python, 622 linespython/ dglke/ models/ mxnet/ score_fun.py - Backend/
dgl-ke/ , Python, 268 linespython/ dglke/ models/ mxnet/ tensor_models.py - Backend/
dgl-ke/ , Python, 18 linespython/ dglke/ models/ pytorch/ __init__.py - Backend/
dgl-ke/ , Python, 248 linespython/ dglke/ models/ pytorch/ ke_tensor.py - Backend/
dgl-ke/ , Python, 98 linespython/ dglke/ models/ pytorch/ loss.py - Backend/
dgl-ke/ , Python, 641 linespython/ dglke/ models/ pytorch/ score_fun.py - Backend/
dgl-ke/ , Python, 407 linespython/ dglke/ models/ pytorch/ tensor_models.py - Backend/
dgl-ke/ , Python, 146 linespython/ dglke/ partition.py - Backend/
dgl-ke/ , Python, 253 linespython/ dglke/ tests/ test_dataset.py - Backend/
dgl-ke/ , Python, 276 linespython/ dglke/ tests/ test_infer.py - Backend/
dgl-ke/ , Python, 214 linespython/ dglke/ tests/ test_score.py - Backend/
dgl-ke/ , Python, 1,264 linespython/ dglke/ tests/ test_topk.py - Backend/
dgl-ke/ , Python, 391 linespython/ dglke/ train.py - Backend/
dgl-ke/ , Python, 133 linespython/ dglke/ train_mxnet.py - Backend/
dgl-ke/ , Python, 403 linespython/ dglke/ train_pytorch.py - Backend/
dgl-ke/ , Python, 350 lines, 2 matchespython/ dglke/ utils.py - Backend/
dgl-ke/ , Python, 25 linespython/ setup.py - Backend/
run_medgemma.sh , Shell, 55 lines - Frontend/
agents.py , Python, 991 lines, 1 match - Frontend/
components/ , Python, 849 lines, 2 matchesAboutUs.py - Frontend/
components/ , Python, 843 linesIntroduction.py - Frontend/
components/ , Python, 2,228 lines, 2 matchesLoginSignup.py - Frontend/
components/ , Python, 790 lines, 1 matchMicroServices.py - Frontend/
components/ , Python, 433 linesUnifiedChat.py - Frontend/
components/ , Python, 149 linesforget_password.py - Frontend/
components/ , Python, 380 lineshow_to_use.py - Frontend/
evo-utils/ , Python, 334 linesdemo_agents.py - Frontend/
evo-utils/ , Python, 75 linesdemo_app.py - Frontend/
evo-utils/ , Python, 8 linessrc/ kani_utils/ __init__.py - Frontend/
evo-utils/ , Python, 86 linessrc/ kani_utils/ base_kanis.py - Frontend/
evo-utils/ , Python, 563 lines, 1 matchsrc/ kani_utils/ kani_streamlit_server.py - Frontend/
evo-utils/ , Python, 808 linessrc/ kani_utils/ kani_streamlit_server_ba ckup.py - Frontend/
evo-utils/ , Python, 53 linessrc/ kani_utils/ utils.py - Frontend/
hypo_agents.py , Python, 712 lines, 2 matches - Frontend/
streamlit_app.py , Python, 499 lines - Frontend/
utils/ , Python, 95 linesaccount_utils.py - Frontend/
utils/ , Python, 183 linesauth_utils.py - pipeline/
01_data_collection/ , Jupyter, 186 linesdata_collection.ipynb - pipeline/
02_data_processing/ , Jupyter, 525 lines1_harmonizome.ipynb - pipeline/
02_data_processing/ , Jupyter, 692 lines2_harmonizome.ipynb - pipeline/
02_data_processing/ , Jupyter, 550 linesAgingbank.ipynb - pipeline/
02_data_processing/ , Jupyter, 820 linesDRKG.ipynb - pipeline/
02_data_processing/ , Jupyter, 207 linesMetaboAge.ipynb - pipeline/
02_data_processing/ , Jupyter, 1,112 lines, 1 matchPrimeKG.ipynb - pipeline/
02_data_processing/ , Jupyter, 356 linesSTITCH_STRING_Human_KG.i pynb - pipeline/
02_data_processing/ , Jupyter, 440 linesSTRING.ipynb - pipeline/
02_data_processing/ , Jupyter, 334 linesageanno.ipynb - pipeline/
02_data_processing/ , Jupyter, 1,212 lines, 1 matchageannomo.ipynb - pipeline/
02_data_processing/ , Jupyter, 629 linesagextend.ipynb - pipeline/
02_data_processing/ , Jupyter, 460 linesaging_atlas.ipynb - pipeline/
02_data_processing/ , Jupyter, 644 linesbiningdb_sms_biosnap_che mbl_drugbank.ipynb - pipeline/
02_data_processing/ , Jupyter, 569 linesbock_kg.ipynb - pipeline/
02_data_processing/ , Jupyter, 1,432 linesckg.ipynb - pipeline/
02_data_processing/ , Jupyter, 806 linescrossbar.ipynb - pipeline/
02_data_processing/ , Jupyter, 539 linesdigital_aging_atlas.ipyn b - pipeline/
02_data_processing/ , Jupyter, 506 lines, 1 matchdrugage.ipynb - pipeline/
02_data_processing/ , Jupyter, 175 linesevolf.ipynb - pipeline/
02_data_processing/ , Jupyter, 452 linesgenage.ipynb - pipeline/
02_data_processing/ , Jupyter, 364 linesgendr.ipynb - pipeline/
02_data_processing/ , Jupyter, 562 lineshald.ipynb - pipeline/
02_data_processing/ , Jupyter, 691 lineshetionet.ipynb - pipeline/
02_data_processing/ , Jupyter, 990 linesibkh.ipynb - pipeline/
02_data_processing/ , Jupyter, 1,138 linesimmpat_ttd_phytohub_othe rsources.ipynb - pipeline/
02_data_processing/ , Jupyter, 212 linesmirTARbase.ipynb - pipeline/
02_data_processing/ , Jupyter, 1,395 linesmonarchkg1.ipynb - pipeline/
02_data_processing/ , Jupyter, 3,203 linesmonarchkg2.ipynb - pipeline/
02_data_processing/ , Jupyter, 538 linespharmkg.ipynb - pipeline/
02_data_processing/ , Jupyter, 988 linesphenknowlator.ipynb - pipeline/
02_data_processing/ , Jupyter, 134 linesspecies_specific_data_so urce/ Celegans/ fic.ipynb - pipeline/
02_data_processing/ , Jupyter, 281 linesspecies_specific_data_so urce/ Celegans/ stitch.ipynb - pipeline/
02_data_processing/ , Jupyter, 251 linesspecies_specific_data_so urce/ Celegans/ string.ipynb - pipeline/
02_data_processing/ , Jupyter, 186 linesspecies_specific_data_so urce/ Celegans/ wormbase.ipynb - pipeline/
02_data_processing/ , Jupyter, 82 linesspecies_specific_data_so urce/ Celegans/ worminteractomedataabse. ipynb - pipeline/
02_data_processing/ , Jupyter, 93 linesspecies_specific_data_so urce/ Drosophila/ flybase.ipynb - pipeline/
02_data_processing/ , Jupyter, 327 linesspecies_specific_data_so urce/ Drosophila/ stitch.ipynb - pipeline/
02_data_processing/ , Jupyter, 274 linesspecies_specific_data_so urce/ Drosophila/ string.ipynb - pipeline/
02_data_processing/ , Jupyter, 238 linesspecies_specific_data_so urce/ Mouse/ mgido.ipynb - pipeline/
02_data_processing/ , Jupyter, 237 linesspecies_specific_data_so urce/ Mouse/ stitch.ipynb - pipeline/
02_data_processing/ , Jupyter, 274 linesspecies_specific_data_so urce/ Mouse/ string_mousenet.ipynb - pipeline/
02_data_processing/ , Jupyter, 147 linesspecies_specific_data_so urce/ Yeast/ esldb.ipynb - pipeline/
02_data_processing/ , Jupyter, 157 linesspecies_specific_data_so urce/ Yeast/ sgd.ipynb - pipeline/
02_data_processing/ , Jupyter, 142 linesspecies_specific_data_so urce/ Yeast/ stitch_ch_ch.ipynb - pipeline/
02_data_processing/ , Jupyter, 143 linesspecies_specific_data_so urce/ Yeast/ stitch_ch_ge.ipynb - pipeline/
02_data_processing/ , Jupyter, 407 linesspecies_specific_data_so urce/ Yeast/ string_yeastnet_biogrid. ipynb - pipeline/
02_data_processing/ , Jupyter, 368 linesspecies_specific_data_so urce/ Zebrafish/ stitch.ipynb - pipeline/
02_data_processing/ , Jupyter, 313 linesspecies_specific_data_so urce/ Zebrafish/ string.ipynb - pipeline/
02_data_processing/ , Jupyter, 477 linesspecies_specific_data_so urce/ Zebrafish/ zfin.ipynb - pipeline/
02_data_processing/ , Jupyter, 1,235 lines, 1 matchtarkg.ipynb - pipeline/
03_relation_merge/ , Jupyter, 171 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Chemic al_Inhibit_BiologicalPro cess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 151 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Chemic al_Noeffect_BiologicalPr ocess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 276 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Chemic al_POS_NEG_NOT_ASSOCIATE D_BiologicalProcess.ipyn b - pipeline/
03_relation_merge/ , Jupyter, 150 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Chemic al_Promotes_BiologicalPr ocess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 167 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_C ellularComponent.ipynb - pipeline/
03_relation_merge/ , Jupyter, 207 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_G ene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 185 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_I nhibit_BiologicalProcess .ipynb - pipeline/
03_relation_merge/ , Jupyter, 188 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_P henotype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 197 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_P romotes_BiologicalProces s.ipynb - pipeline/
03_relation_merge/ , Jupyter, 160 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_a natomy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 193 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_b iologicalprocess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 171 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Gene_m olecularfunction.ipynb - pipeline/
03_relation_merge/ , Jupyter, 157 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Phenot ype_biologicalprocess.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 151 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Phenot ype_chemicalentity.ipynb - pipeline/
03_relation_merge/ , Jupyter, 157 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Phenot ype_molecularfunction.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 169 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_Protei n_Tissue.ipynb - pipeline/
03_relation_merge/ , Jupyter, 154 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_anatom y_anatomy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 172 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_chemic al_gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 148 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_chemic al_protein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 172 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_gene_p athway.ipynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_mirna_ gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 167 linesOTHER_SPECIES/ Celegans/ Celegans_Rel_wise_phenot ype_CellularComponent.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 171 linesOTHER_SPECIES/ Drosophila/ Droso_Rel_wise_Chemical_ Inhibit_BiologicalProces s.ipynb - pipeline/
03_relation_merge/ , Jupyter, 313 linesOTHER_SPECIES/ Drosophila/ Droso_Rel_wise_Chemical_ POS_NEG_NOT_ASSOCIATED_B iologicalProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 150 linesOTHER_SPECIES/ Drosophila/ Droso_Rel_wise_Chemical_ Promotes_BiologicalProce ss.ipynb - pipeline/
03_relation_merge/ , Jupyter, 177 linesOTHER_SPECIES/ Drosophila/ Droso_Rel_wise_Gene_Inhi bit_BiologicalProcess.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 197 linesOTHER_SPECIES/ Drosophila/ Droso_Rel_wise_Gene_Prom otes_BiologicalProcess.i pynb - pipeline/
03_relation_merge/ , Jupyter, 195 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_Gene _Gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 203 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_Gene _biologicalproces.ipynb - pipeline/
03_relation_merge/ , Jupyter, 159 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_Prot ein_Tissue.ipynb - pipeline/
03_relation_merge/ , Jupyter, 171 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_anat omy_anatomy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_anat omy_cellualrcomponent.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 176 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_chem ical_gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 164 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_chem ical_protein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 168 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_gene _MolecularFunction.ipynb - pipeline/
03_relation_merge/ , Jupyter, 176 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_gene _anatomy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 168 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_gene _cellularcomponent.ipynb - pipeline/
03_relation_merge/ , Jupyter, 168 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_gene _pathway.ipynb - pipeline/
03_relation_merge/ , Jupyter, 186 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_gene _phenotype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 174 linesOTHER_SPECIES/ Drosophila/ Drosophila_Rel_wise_mirn a_gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 269 linesOTHER_SPECIES/ Mouse/ Mouse_Rel_wise_Chemical_ POS_NEG_NOT_ASSOCIATED_B iologicalProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 187 linesOTHER_SPECIES/ Mouse/ Mouse_Rel_wise_chemical_ gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Chemical_ BiologicalProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 155 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Chemical_ Chemical.ipynb - pipeline/
03_relation_merge/ , Jupyter, 168 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Chemical_ Inhibit_BiologicalProces s.ipynb - pipeline/
03_relation_merge/ , Jupyter, 209 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_Gene .ipynb - pipeline/
03_relation_merge/ , Jupyter, 204 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_Inhi bit_BiologicalProcess.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 254 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_PosA ssociation_NegAssociatio n_NoAssociation_Biologic alProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 210 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_Prom otes_BiologicalProcess.i pynb - pipeline/
03_relation_merge/ , Jupyter, 214 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_biol ogical_process.ipynb - pipeline/
03_relation_merge/ , Jupyter, 168 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_cell ularcomponent.ipynb - pipeline/
03_relation_merge/ , Jupyter, 174 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_mole cularfunction.ipynb - pipeline/
03_relation_merge/ , Jupyter, 299 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_phen otype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 165 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_Gene_tiss ue.ipynb - pipeline/
03_relation_merge/ , Jupyter, 182 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_anatomy_a natomy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 172 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_mirna_gen e.ipynb - pipeline/
03_relation_merge/ , Jupyter, 179 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_phenotype _biologicalprocess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 167 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_phenotype _cellularcomponent.ipynb - pipeline/
03_relation_merge/ , Jupyter, 167 linesOTHER_SPECIES/ Mouse/ mouse_Rel_wise_phenotype _chemical-Copy1.ipynb - pipeline/
03_relation_merge/ , Jupyter, 184 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Chemical_ Inhibit_BiologicalProces s.ipynb - pipeline/
03_relation_merge/ , Jupyter, 273 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Chemical_ POS_NEG_NOT_ASSOCIATED_B iologicalProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 142 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Gene_Anat omy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 156 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Gene_Cell ularComponent.ipynb - pipeline/
03_relation_merge/ , Jupyter, 173 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Gene_Gene .ipynb - pipeline/
03_relation_merge/ , Jupyter, 177 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Gene_Inhi bit_BiologicalProcess.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 143 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Gene_Mole cularFunction.ipynb - pipeline/
03_relation_merge/ , Jupyter, 155 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Gene_Phen otype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 181 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_Gene_Prom otes_BiologicalProcess.i pynb - pipeline/
03_relation_merge/ , Jupyter, 189 linesOTHER_SPECIES/ Yeast/ Yeast_Rel_wise_chemical_ gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 182 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_Gene_ BiologicalProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 234 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_Gene_ Gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 248 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_Gene_ Phenotype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 163 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_Gene_ anatomy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 173 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_Gene_ cellularcomponent.ipynb - pipeline/
03_relation_merge/ , Jupyter, 167 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_Gene_ molecularfunction.ipynb - pipeline/
03_relation_merge/ , Jupyter, 167 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_Gene_ pathway.ipynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_anato my_anatomy.ipynb - pipeline/
03_relation_merge/ , Jupyter, 171 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_chemi cal_gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 168 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_chemi cal_protein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 169 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_mirna _gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_pheno type_biologicalprocess.i pynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_pheno type_cellularcomponent.i pynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_pheno type_chemicalentity.ipyn b - pipeline/
03_relation_merge/ , Jupyter, 169 linesOTHER_SPECIES/ Zebrafish/ zebrafish_Rel_wise_pheno type_molecularfunction.i pynb - pipeline/
03_relation_merge/ , Jupyter, 246 linesRelation_Wise_ANATOMY_AN ATOMY.ipynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesRelation_Wise_ANATOMY_Bi ologicalProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 158 linesRelation_Wise_ANATOMY_CH EMICAl.ipynb - pipeline/
03_relation_merge/ , Jupyter, 297 linesRelation_Wise_ANATOMY_GE NE.ipynb - pipeline/
03_relation_merge/ , Jupyter, 170 linesRelation_Wise_ANATOMY_ce llularComp.ipynb - pipeline/
03_relation_merge/ , Jupyter, 296 linesRelation_Wise_BioloProce ss_Chemical.ipynb - pipeline/
03_relation_merge/ , Jupyter, 249 linesRelation_Wise_Biological Pro_Gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 288 linesRelation_Wise_Biological Process_BiologicalProces s.ipynb - pipeline/
03_relation_merge/ , Jupyter, 171 linesRelation_Wise_Biological Process_Cellularcomnp.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 157 linesRelation_Wise_Biological Process_MolecularFunctio n.ipynb - pipeline/
03_relation_merge/ , Jupyter, 157 linesRelation_Wise_Biological Process_Protein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 158 linesRelation_Wise_CHEMICAL_I NHIBITS_BIOLOGICALPROCES S_Human.ipynb - pipeline/
03_relation_merge/ , Jupyter, 126 linesRelation_Wise_CHEMICAL_N oeffect_BIOLOGICALPROCES S_Human.ipynb - pipeline/
03_relation_merge/ , Jupyter, 155 linesRelation_Wise_CHEMICAL_P OS_NEG_ASSOCIATION_BIOLO GICALPROCESS_Human.ipynb - pipeline/
03_relation_merge/ , Jupyter, 138 linesRelation_Wise_CHEMICAL_P ROMOTES_BIOLOGICALPROCES S_Human.ipynb - pipeline/
03_relation_merge/ , Jupyter, 230 linesRelation_Wise_CellularCo mponent_CellularComponen t.ipynb - pipeline/
03_relation_merge/ , Jupyter, 139 linesRelation_Wise_CellularCo mponent__ChemicalENTITY. ipynb - pipeline/
03_relation_merge/ , Jupyter, 188 linesRelation_Wise_ChemicalEn tity_BiologicalProcess.i pynb - pipeline/
03_relation_merge/ , Jupyter, 512 linesRelation_Wise_Chemical_C hemical.ipynb - pipeline/
03_relation_merge/ , Jupyter, 503 linesRelation_Wise_Disease_An atomy_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 592 linesRelation_Wise_Disease_Ch emical.ipynb - pipeline/
03_relation_merge/ , Jupyter, 512 linesRelation_Wise_Disease_Di sease.ipynb - pipeline/
03_relation_merge/ , Jupyter, 519 linesRelation_Wise_Disease_Ge ne.ipynb - pipeline/
03_relation_merge/ , Jupyter, 299 linesRelation_Wise_Disease_Mu tation.ipynb - pipeline/
03_relation_merge/ , Jupyter, 369 linesRelation_Wise_Disease_Pa thway.ipynb - pipeline/
03_relation_merge/ , Jupyter, 417 linesRelation_Wise_Disease_Ph enotype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 190 linesRelation_Wise_GENE_ANATO MY.ipynb - pipeline/
03_relation_merge/ , Jupyter, 407 linesRelation_Wise_GENE_BIOLO GICALPROCESS.ipynb - pipeline/
03_relation_merge/ , Jupyter, 256 linesRelation_Wise_GENE_CELLU LARCOMPONENT.ipynb - pipeline/
03_relation_merge/ , Jupyter, 365 linesRelation_Wise_GENE_CHEMI CAL_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 361 linesRelation_Wise_GENE_GENE. ipynb - pipeline/
03_relation_merge/ , Jupyter, 292 linesRelation_Wise_GENE_PHENO TYPE_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 145 linesRelation_Wise_GENE_POS_N EG_NO_ASSOCIATED_BIOLOGI CALPROCESS_Human.ipynb - pipeline/
03_relation_merge/ , Jupyter, 141 linesRelation_Wise_GENE_PROMO TES_BIOLOFICALPROCESS-Co py1.ipynb - pipeline/
03_relation_merge/ , Jupyter, 139 linesRelation_Wise_GENE_PROMO TES_BIOLOFICALPROCESS.ip ynb - pipeline/
03_relation_merge/ , Jupyter, 224 linesRelation_Wise_GENE_PROTE IN.ipynb - pipeline/
03_relation_merge/ , Jupyter, 574 linesRelation_Wise_Gene_Disea se.ipynb - pipeline/
03_relation_merge/ , Jupyter, 293 linesRelation_Wise_Gene_Molec ularFunction_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 150 linesRelation_Wise_Gene_Mutat ion.ipynb - pipeline/
03_relation_merge/ , Jupyter, 300 linesRelation_Wise_Gene_Pathw ay.ipynb - pipeline/
03_relation_merge/ , Jupyter, 198 linesRelation_Wise_Gene_Tissu e.ipynb - pipeline/
03_relation_merge/ , Jupyter, 157 linesRelation_Wise_Gene_inhib it_BioProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 187 linesRelation_Wise_Mirna_Gene .ipynb - pipeline/
03_relation_merge/ , Jupyter, 166 linesRelation_Wise_MoleFuncti on_ChemicalENtity.ipynb - pipeline/
03_relation_merge/ , Jupyter, 258 linesRelation_Wise_MoleFuncti on_MoleFunction.ipynb - pipeline/
03_relation_merge/ , Jupyter, 246 linesRelation_Wise_MoleFuncti on_MoleFunction_old.ipyn b - pipeline/
03_relation_merge/ , Jupyter, 95 linesRelation_Wise_MolecularF unction_BiologicalProces s.ipynb - pipeline/
03_relation_merge/ , Jupyter, 134 linesRelation_Wise_MolecularF unction_Protein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 357 linesRelation_Wise_Mutaion_Di sease.ipynb - pipeline/
03_relation_merge/ , Jupyter, 167 linesRelation_Wise_Mutation_C hemical.ipynb - pipeline/
03_relation_merge/ , Jupyter, 226 linesRelation_Wise_Mutation_G ene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 229 linesRelation_Wise_Mutation_G ene_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 146 linesRelation_Wise_Mutation_M utation.ipynb - pipeline/
03_relation_merge/ , Jupyter, 222 linesRelation_Wise_Mutation_P rotein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 168 linesRelation_Wise_PATHWAY_Ge ne.ipynb - pipeline/
03_relation_merge/ , Jupyter, 121 linesRelation_Wise_PMID_Cellu larComp.ipynb - pipeline/
03_relation_merge/ , Jupyter, 105 linesRelation_Wise_PMID_Chemi cal.ipynb - pipeline/
03_relation_merge/ , Jupyter, 109 linesRelation_Wise_PMID_Disea se.ipynb - pipeline/
03_relation_merge/ , Jupyter, 124 linesRelation_Wise_PMID_Prote in.ipynb - pipeline/
03_relation_merge/ , Jupyter, 107 linesRelation_Wise_PMID_Tissu e.ipynb - pipeline/
03_relation_merge/ , Jupyter, 160 linesRelation_Wise_PROTEIN_CH EMICAL.ipynb - pipeline/
03_relation_merge/ , Jupyter, 466 linesRelation_Wise_PROTEIN_Di sease.ipynb - pipeline/
03_relation_merge/ , Jupyter, 238 linesRelation_Wise_PROTEIN_MO LECULARFUNCTION.ipynb - pipeline/
03_relation_merge/ , Jupyter, 237 linesRelation_Wise_PROTEIN_Pa thway_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 175 linesRelation_Wise_Pathway_Pa thway_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 178 linesRelation_Wise_Phenotype_ ChemicalEntity.ipynb - pipeline/
03_relation_merge/ , Jupyter, 304 linesRelation_Wise_Phenotype_ Disease.ipynb - pipeline/
03_relation_merge/ , Jupyter, 162 linesRelation_Wise_Phenotype_ Gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 226 linesRelation_Wise_Phenotype_ Phenotype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 174 linesRelation_Wise_Phenotype_ Protein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 244 linesRelation_Wise_Protein_Bi ologicalProcess.ipynb - pipeline/
03_relation_merge/ , Jupyter, 241 linesRelation_Wise_Protein_Ce llularComponent.ipynb - pipeline/
03_relation_merge/ , Jupyter, 142 linesRelation_Wise_Protein_Ph enotype.ipynb - pipeline/
03_relation_merge/ , Jupyter, 223 linesRelation_Wise_Protein_Ti ssue_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 245 linesRelation_wise_Cellularco mponent_gene.ipynb - pipeline/
03_relation_merge/ , Jupyter, 821 linesRelation_wise_ChemicalEn tity_Disease.ipynb - pipeline/
03_relation_merge/ , Jupyter, 528 linesRelation_wise_ChemicalEn tity_GENE.ipynb - pipeline/
03_relation_merge/ , Jupyter, 320 linesRelation_wise_ChemicalEn tity_Mutation.ipynb - pipeline/
03_relation_merge/ , Jupyter, 279 linesRelation_wise_ChemicalEn tity_Pathway.ipynb - pipeline/
03_relation_merge/ , Jupyter, 433 linesRelation_wise_ChemicalEn tity_Protein.ipynb - pipeline/
03_relation_merge/ , Jupyter, 175 linesRelation_wise_ChemicalEn tity_Tissue.ipynb - pipeline/
03_relation_merge/ , Jupyter, 300 linesRelation_wise_PROTEIN_PR OTEIN_clean.ipynb - pipeline/
03_relation_merge/ , Jupyter, 419 linesRelation_wise_Plant_Dise ase.ipynb - pipeline/
03_relation_merge/ , Jupyter, 306 linesRelation_wise_Plant_chem ical.ipynb - pipeline/
03_relation_merge/ , Python, 191 linesRun1_Standardize_labels. py - pipeline/
03_relation_merge/ , Python, 122 linesRun2_check_humanKG_files and_kg_type_and_add_spec ies_col.py - pipeline/
04_orthology_mapping/ , Python, 125 linesRun_1_collecting_all_spe cies_Unique_gene.py - pipeline/
04_orthology_mapping/ , R, 350 lines, 2 matchesRun_2_Final_map_ortholog s_to_human_desc.R - pipeline/
04_orthology_mapping/ , Python, 227 linesRun_3_1to1_converting.py - pipeline/
04_orthology_mapping/ , Python, 42 linesRun_4_merge_121_12M_toma ke_121PLUS12M.py - pipeline/
04_orthology_mapping/ , Python, 288 linesRun_5_1toM_121_convertin g.py - pipeline/
05_kg_construction/ , Python, 119 linesRun1_split_kg_by_type_hu man.py - pipeline/
05_kg_construction/ , Python, 147 linesRun2_split_1to1_kg_by_ty pe_otherspecies.py - pipeline/
05_kg_construction/ , Python, 163 linesRun3_split_12M_121_comb_ kg_by_type_otherspecies. py - pipeline/
05_kg_construction/ , Python, 168 linesRun_4_making_species_ass ociatedwith_connection/ Run_4.1_1_to_1_ortholog/ Run_4.1_make_species_tri ples_1_to_1.py - pipeline/
05_kg_construction/ , Python, 182 linesRun_4_making_species_ass ociatedwith_connection/ Run_4.2_one2one_plus_one 2many_ortholog/ Run_4.2_make_species_tri ples_combined121_12M.py - pipeline/
06_tensor_building/ , Python, 500 lines01.py - pipeline/
06_tensor_building/ , Python, 314 linesbuilding_aging_kg_new/ Run_1_Aging_1_to_1.py - pipeline/
06_tensor_building/ , Python, 402 linesbuilding_aging_kg_new/ Run_2_Aging_121_12M.py - pipeline/
06_tensor_building/ , Python, 374 linesbuilding_biomedical_kg_n ew/ Run_1_Biomedical_1_to_1. py - pipeline/
06_tensor_building/ , Python, 403 linesbuilding_biomedical_kg_n ew/ Run_2_Biomedical_121_12M .py - pipeline/
06_tensor_building/ , Python, 724 linesbuilding_evoage_kg_new/ Run_1_EvoAge_1_to_1.py - pipeline/
06_tensor_building/ , Python, 183 linesbuilding_evoage_kg_new/ Run_2_EvoAge_121_12M.py - pipeline/
06_tensor_building/ , Python, 500 linesmaking_global_mapping.py - pipeline/
07_model_training/ , Shell, 33 linesAging_121_12M/ Rescal_64.sh - pipeline/
07_model_training/ , Shell, 33 linesAging_1_to_1/ Rescal_64.sh - pipeline/
07_model_training/ , Shell, 33 linesBiomedical_121_12M/ Rescal_64.sh - pipeline/
07_model_training/ , Shell, 25 linesBiomedical_121_12M/ Rescal_64_Biomed_121_12M _testing.sh - pipeline/
07_model_training/ , Python, 165 linesBiomedical_121_12M/ edge_type/ B_121_12M_64_rescal_edge _type_metrics.py - pipeline/
07_model_training/ , Shell, 34 linesBiomedical_1_to_1/ Rescal_64.sh - pipeline/
07_model_training/ , Shell, 26 linesBiomedical_1_to_1/ Rescal_64_Biomed_1_to_1_ testing.sh - pipeline/
07_model_training/ , Python, 165 linesBiomedical_1_to_1/ edge_type/ B_121_64_rescal_edge_typ e_metrics.py - pipeline/
07_model_training/ , Python, 189 linesEvoAge_121_12M/ Run2_64_edge_type_testin g/ 64_complex_edge_type_met ric.py - pipeline/
07_model_training/ , Python, 150 linesEvoAge_121_12M/ Run2_64_edge_type_testin g/ 64_dismult_edge_type_met rics.py - pipeline/
07_model_training/ , Python, 158 linesEvoAge_121_12M/ Run2_64_edge_type_testin g/ 64_rescal_edge_type_metr ics.py - pipeline/
07_model_training/ , Python, 175 linesEvoAge_121_12M/ Run2_64_edge_type_testin g/ 64_rotate_edge_type_metr ics.py - pipeline/
07_model_training/ , Python, 173 linesEvoAge_121_12M/ Run2_64_edge_type_testin g/ 64_simple_edge_type_metr ics.py - pipeline/
07_model_training/ , Python, 162 linesEvoAge_121_12M/ Run2_64_edge_type_testin g/ 64_transE_edgetype_metri cs.py - pipeline/
07_model_training/ , Shell, 29 linesEvoAge_121_12M/ Run_1_train_test/ Complex_64_testing.sh - pipeline/
07_model_training/ , Shell, 34 linesEvoAge_121_12M/ Run_1_train_test/ Dismult.sh - pipeline/
07_model_training/ , Shell, 29 linesEvoAge_121_12M/ Run_1_train_test/ Dismult_64_testing.sh - pipeline/
07_model_training/ , Shell, 33 linesEvoAge_121_12M/ Run_1_train_test/ Rescal_64.sh - pipeline/
07_model_training/ , Shell, 29 linesEvoAge_121_12M/ Run_1_train_test/ Rescal_64_testing.sh - pipeline/
07_model_training/ , Shell, 36 linesEvoAge_121_12M/ Run_1_train_test/ RotatE_64.sh - pipeline/
07_model_training/ , Shell, 30 linesEvoAge_121_12M/ Run_1_train_test/ RotatE_64_testing.sh - pipeline/
07_model_training/ , Shell, 36 linesEvoAge_121_12M/ Run_1_train_test/ Simple_64.sh - pipeline/
07_model_training/ , Shell, 29 linesEvoAge_121_12M/ Run_1_train_test/ Simple_64_testing.sh - pipeline/
07_model_training/ , Shell, 29 linesEvoAge_121_12M/ Run_1_train_test/ Transe_64_testing.sh - pipeline/
07_model_training/ , Shell, 32 linesEvoAge_121_12M/ Run_1_train_test/ complex_64.sh - pipeline/
07_model_training/ , Shell, 34 linesEvoAge_121_12M/ Run_1_train_test/ transe_64.sh - pipeline/
07_model_training/ , Shell, 34 linesEvoAge_1_to_1/ Rescal_64.sh - pipeline/
07_model_training/ , Shell, 29 linesEvoAge_1_to_1/ Rescal_64_Evoage_1_to_1_ testing.sh - pipeline/
08_evaluation_analysis/ , Python, 64 linesAgingonly_testing_data/ Aging_testing_data_on_al lKG_compling.py - pipeline/
08_evaluation_analysis/ , Shell, 63 linesAgingonly_testing_data/ Aging_testing_data_on_al lKGs.sh - pipeline/
08_evaluation_analysis/ , Jupyter, 773 linesbuilding_evoage_with_1_p ercent_species_testsplit / Training/ Run1_percent_species_tes t_set.ipynb - pipeline/
08_evaluation_analysis/ , Shell, 34 linesbuilding_evoage_with_1_p ercent_species_testsplit / Training/ Run2_per_Rescal_64.sh - pipeline/
08_evaluation_analysis/ , Shell, 121 linesbuilding_evoage_with_1_p ercent_species_testsplit / Training/ Run3_Rescal_64_testing.s h - pipeline/
08_evaluation_analysis/ , Python, 217 linesbuilding_evoage_with_1_p ercent_species_testsplit / Training/ Run4_per_rescal_edge_typ e_metrics.py - pipeline/
08_evaluation_analysis/ , Python, 147 linesshuffled_kg_testset/ Run1_make_shuffled_KG.py - pipeline/
08_evaluation_analysis/ , Shell, 48 linesshuffled_kg_testset/ Run2_Testing/ run_rescal_eval_all_shuf fles.sh - pipeline/
08_evaluation_analysis/ , Python, 55 linesshuffled_kg_testset/ Run3_testing_log_to_csv. py - pipeline/
09_evoage_vs_other/ , Python, 303 linesevoage_vs_biochat_escarg ot/ Biochatter/ evoage_biochatter.py - pipeline/
09_evoage_vs_other/ , Python, 141 lines, 1 matchevoage_vs_biochat_escarg ot/ Biochatter/ generate_schema_info.py - pipeline/
09_evoage_vs_other/ , Python, 360 linesevoage_vs_biochat_escarg ot/ Biochatter/ hypothesis_ques_answer/ medgemma_analyse_biochat ter_response.py - pipeline/
09_evoage_vs_other/ , Python, 201 linesevoage_vs_biochat_escarg ot/ Biochatter/ hypothesis_ques_answer/ run_hypothesis_benchmark .py - pipeline/
09_evoage_vs_other/ , Python, 157 linesevoage_vs_biochat_escarg ot/ Biochatter/ run_benchmark.py - pipeline/
09_evoage_vs_other/ , Python, 152 linesevoage_vs_biochat_escarg ot/ analysis/ analyse_verdicts_alterna tive_plots.py - pipeline/
09_evoage_vs_other/ , Python, 73 linesevoage_vs_biochat_escarg ot/ analysis/ analysis_only_verdict_by _label.py - pipeline/
09_evoage_vs_other/ , Python, 320 linesevoage_vs_biochat_escarg ot/ analysis/ verdict_analysis_complet e.py - pipeline/
09_evoage_vs_other/ , Python, 80 linesevoage_vs_biochat_escarg ot/ escargot/ alzkb_sanity/ ask_alzkb.py - pipeline/
09_evoage_vs_other/ , Python, 143 linesevoage_vs_biochat_escarg ot/ escargot/ alzkb_sanity/ load_alzkb_sample.py - pipeline/
09_evoage_vs_other/ , Python, 2 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 1 lineevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ coder/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 115 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ coder/ coder.py - pipeline/
09_evoage_vs_other/ , Python, 1 lineevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ controller/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 315 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ controller/ controller.py - pipeline/
09_evoage_vs_other/ , Python, 2 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ cypher/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 155 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ cypher/ memgraph.py - pipeline/
09_evoage_vs_other/ , Python, 155 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ cypher/ neo4j.py - pipeline/
09_evoage_vs_other/ , Python, 418 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ escargot.py - pipeline/
09_evoage_vs_other/ , Python, 4 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ language_models/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 95 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ language_models/ abstract_language_model. py - pipeline/
09_evoage_vs_other/ , Python, 145 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ language_models/ azuregpt.py - pipeline/
09_evoage_vs_other/ , Python, 166 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ language_models/ chatgpt.py - pipeline/
09_evoage_vs_other/ , Python, 141 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ language_models/ ollama.py - pipeline/
09_evoage_vs_other/ , Python, 1 lineevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ memory/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 135 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ memory/ memory.py - pipeline/
09_evoage_vs_other/ , Python, 1 lineevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ multiagent/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 150 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ multiagent/ multiagent_manager.py - pipeline/
09_evoage_vs_other/ , Python, 334 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ multiagent/ utils.py - pipeline/
09_evoage_vs_other/ , Python, 6 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ operations/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 69 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ operations/ graph_of_operations.py - pipeline/
09_evoage_vs_other/ , Python, 390 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ operations/ operations.py - pipeline/
09_evoage_vs_other/ , Python, 21 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ operations/ thought.py - pipeline/
09_evoage_vs_other/ , Python, 51 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ operations/ utils.py - pipeline/
09_evoage_vs_other/ , Python, 1 lineevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ parser/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 131 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ parser/ parser.py - pipeline/
09_evoage_vs_other/ , Python, 169 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ parser/ utils.py - pipeline/
09_evoage_vs_other/ , Python, 1 lineevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ prompter/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 557 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ prompter/ prompter.py - pipeline/
09_evoage_vs_other/ , Python, 45 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ utils.py - pipeline/
09_evoage_vs_other/ , Python, 2 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ vector_db/ __init__.py - pipeline/
09_evoage_vs_other/ , Python, 14 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ vector_db/ azure_embedding.py - pipeline/
09_evoage_vs_other/ , Python, 143 linesevoage_vs_biochat_escarg ot/ escargot/ escargot/ escargot/ vector_db/ weaviate.py - pipeline/
09_evoage_vs_other/ , Python, 219 linesevoage_vs_biochat_escarg ot/ escargot/ evoage/ ask_evoage.py - pipeline/
09_evoage_vs_other/ , Python, 169 linesevoage_vs_biochat_escarg ot/ escargot/ evoage/ deepseek_lm.py - pipeline/
09_evoage_vs_other/ , Python, 183 linesevoage_vs_biochat_escarg ot/ escargot/ evoage/ run_benchmark.py - pipeline/
09_evoage_vs_other/ , Python, 206 linesevoage_vs_biochat_escarg ot/ escargot/ hypothesis_ques_answer/ medgemma_analyse_escargo t_response.py - pipeline/
09_evoage_vs_other/ , Python, 178 linesevoage_vs_biochat_escarg ot/ escargot/ hypothesis_ques_answer/ run_hypothesis_benchmark .py - pipeline/
09_evoage_vs_other/ , Shell, 94 linesevoage_vs_medgemma_biomi stral/ BioMistral/ run.sh - pipeline/
09_evoage_vs_other/ , Python, 168 linesevoage_vs_medgemma_biomi stral/ BioMistral/ run_hypothesis.py - pipeline/
09_evoage_vs_other/ , Shell, 94 linesevoage_vs_medgemma_biomi stral/ BioMistralFinetuned/ run.sh - pipeline/
09_evoage_vs_other/ , Python, 168 linesevoage_vs_medgemma_biomi stral/ BioMistralFinetuned/ run_hypothesis.py - pipeline/
09_evoage_vs_other/ , Shell, 94 linesevoage_vs_medgemma_biomi stral/ medgemma/ run.sh - pipeline/
09_evoage_vs_other/ , Python, 168 linesevoage_vs_medgemma_biomi stral/ medgemma/ run_hypothesis.py - pipeline/
09_evoage_vs_other/ , Shell, 41 linesevoage_vs_medgemma_biomi stral/ run_all.sh - pipeline/
10_llm_fintunning/ , Python, 301 linesload_kuzu.py - pipeline/
10_llm_fintunning/ , Python, 180 linesload_kuzu_test.py - pipeline/
10_llm_fintunning/ , Shell, 132 linespipeline_data_gen.sh - pipeline/
10_llm_fintunning/ , Shell, 176 linespipeline_eval.sh - pipeline/
10_llm_fintunning/ , Shell, 80 linespipeline_sft.sh - pipeline/
10_llm_fintunning/ , Python, 143 linesregister_datasets.py - pipeline/
10_llm_fintunning/ , Python, 256 linesresolve_kg_evaluation.py - pipeline/
10_llm_fintunning/ , Shell, 136 linesrun_pipeline.sh - pipeline/
11_ploting/ , Jupyter, 5,312 lines, 2 matchesall_figures-raidhani.ipy nb - repository limit reached (2,000 files or 30 MB): the rest is at the source (2 files)
- LICENSE, License, 21 lines
- README.md, Text, 588 lines
hub.docker.com/r/ahujalab
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
- 30 September 2026: the link answers (HTTP 200)
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 419 scripts, each with its path and the digest of its content;
- 21 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- zenodo:17711173, at Zenodo; found in “Data and Code Availability”
Data and Code Availability
The complete source code for the EvoAge platform is publicly available on GitHub (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 30 September 2026: the first record
Recorded: type, journal, dates, 21 authors, 6 keywords, 84 references.
Cite
This paper
Ahuja, G., Sharma, A., Singh, A., Kedia, S., Sharma, A., Gautam, V., Sharma, P., Khandelwal, A., Ramakrishna, S., Muddashetty, R., Freude, K., Solanki, S., Chauhan, S., Kumar, S., Satija, S., Duari, S., Arora, S., Gupta, A., Shome, R., . . . nair, D. (2026). Cross-Species Aging Knowledge Integration into Agentic AI Platform Uncovers Conserved Mechanisms. Research Square (preprint). https://
BibTeX
@article{ahuja2026cross,
author = {Ahuja, Gaurav and Sharma, Arushi and Singh, Ankit and Kedia, Shekhar and Sharma, Abhinav and Gautam, Vishakha and Sharma, Pranjal and Khandelwal, Aniket and Ramakrishna, Sarayu and Muddashetty, Ravi and Freude, Kristine and Solanki, Saveena and Chauhan, Sonam and Kumar, Suvendu and Satija, Shiva and Duari, Subhadeep and Arora, Sakshi and Gupta, Advik and Shome, Raidhani and Sengupta, Debarka and nair, Deepak},
title = {{Cross-Species Aging Knowledge Integration into Agentic AI Platform Uncovers Conserved Mechanisms}},
journal = {Research Square (preprint)},
year = {2026},
month = mar,
publisher = {Research Square},
issn = {2693-5015},
doi = {10.21203/
url = {https://
}
RIS
TY - JOUR
AU - Ahuja, Gaurav
AU - Sharma, Arushi
AU - Singh, Ankit
AU - Kedia, Shekhar
AU - Sharma, Abhinav
AU - Gautam, Vishakha
AU - Sharma, Pranjal
AU - Khandelwal, Aniket
AU - Ramakrishna, Sarayu
AU - Muddashetty, Ravi
AU - Freude, Kristine
AU - Solanki, Saveena
AU - Chauhan, Sonam
AU - Kumar, Suvendu
AU - Satija, Shiva
AU - Duari, Subhadeep
AU - Arora, Sakshi
AU - Gupta, Advik
AU - Shome, Raidhani
AU - Sengupta, Debarka
AU - nair, Deepak
TI - Cross-Species Aging Knowledge Integration into Agentic AI Platform Uncovers Conserved Mechanisms
T2 - Research Square (preprint)
J2 - Res Sq
PY - 2026
DA - 2026/
SN - 2693-5015
PB - Research Square
DO - 10.21203/
UR - https://
ER -
CSL-JSON
{
"id": "10.21203/
"type": "article",
"title": "Cross-Species Aging Knowledge Integration into Agentic AI Platform Uncovers Conserved Mechanisms",
"container-title": "Research Square (preprint)",
"author": [
{
"family": "Ahuja",
"given": "Gaurav"
},
{
"family": "Sharma",
"given": "Arushi"
},
{
"family": "Singh",
"given": "Ankit"
},
{
"family": "Kedia",
"given": "Shekhar"
},
{
"family": "Sharma",
"given": "Abhinav"
},
{
"family": "Gautam",
"given": "Vishakha"
},
{
"family": "Sharma",
"given": "Pranjal"
},
{
"family": "Khandelwal",
"given": "Aniket"
},
{
"family": "Ramakrishna",
"given": "Sarayu"
},
{
"family": "Muddashetty",
"given": "Ravi"
},
{
"family": "Freude",
"given": "Kristine"
},
{
"family": "Solanki",
"given": "Saveena"
},
{
"family": "Chauhan",
"given": "Sonam"
},
{
"family": "Kumar",
"given": "Suvendu"
},
{
"family": "Satija",
"given": "Shiva"
},
{
"family": "Duari",
"given": "Subhadeep"
},
{
"family": "Arora",
"given": "Sakshi"
},
{
"family": "Gupta",
"given": "Advik"
},
{
"family": "Shome",
"given": "Raidhani"
},
{
"family": "Sengupta",
"given": "Debarka"
},
{
"family": "nair",
"given": "Deepak"
}
],
"container-title-short":
"DOI": "10.21203/
"ISSN": "2693-5015",
"publisher": "Research Square",
"URL": "https://
"issued": {
"date-parts": [
[
2026,
3,
13
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: CuPy, NetworkX, Plotly, 8 other tools, cellular / molecular
- [2] doi:10.1016/j.isci.2026.116825 [code]
- Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.Journal: iScienceIn common: CuPy, NetworkX, Plotly, 7 other tools
- [3] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: CuPy, NetworkX, Plotly, 7 other tools
- [4] doi:10.1162/imag.a.1276 [code]
- High-resolution whole-brain magnetic resonance spectroscopic imaging in youth at risk for psychosis.Journal: Imaging neuroscience (Cambridge, Mass.)In common: CuPy, NetworkX, Plotly, 6 other tools
- [5] doi:10.1186/s13293-026-00927-4 [code]
- Gene regulatory network analysis identifies dysregulation of hypoxia pathways as contributing to glioblastoma treatment resistance in females.Journal: Biology of sex differencesIn common: CuPy, NetworkX, PyTorch, 5 other tools, cellular / molecular, 1 reference
- [6] doi:10.1038/s41514-026-00391-9 [code]
- Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.Journal: npj agingIn common: Plotly, PyTorch, tidyverse, 5 other tools, cellular / molecular, 2 references
- [7] doi:10.1371/journal.pcbi.1014555 [code]
- Body surface potential driven personalisation of electrophysiological digital twins in hypertrophic cardiomyopathy.Journal: PLoS computational biologyIn common: CuPy, NetworkX, Pillow, 6 other tools
- [8] doi:10.64898/2026.03.30.714220 [code]
- An integrated single cell and spatial omics atlas of human prenatal developmentJournal: bioRxiv (preprint)In common: CuPy, NetworkX, Pillow, 6 other tools
- [9] doi:10.1038/s41593-026-02267-3 [code]
- Spatial proteomic analysis in human Alzheimer's disease brains enables identification of microenvironment-depende
nt microglial cell states. Journal: Nature neuroscienceIn common: CuPy, NetworkX, Plotly, 5 other tools, Alzheimer's / dementia, cellular / molecular - [10] doi:10.1038/s41467-026-71525-6 [code]
- Single-nucleus brain transcriptomics reveals microglia dysfunction in multiple system atrophy.Journal: Nature communicationsIn common: tidyverse, pandas, Matplotlib, 1 other tool, cellular / molecular, 1 reference, author Kristine Freude
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 419 scripts, and 21 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:7b17ab7a38413b83…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
