OSCR

A microprotein atlas of the human frontal cortex in Alzheimer's disease.

Code ↔ Paper

29 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 29 matches
  1. [1] § Methods › Spectral validation of MP PSMs using PROSIT ↔ Code/Peptide_TMT_analysis/prosit/prosit_pipeline.py, lines 726–867 · score 0.95 · matched fragment ions, PROSIT predicted fragments, lowest SA, SA rating, matched predicted, Match coverage
  2. [2] § Methods › Cell culture, transient transfection and immunocytochemistry ↔ Code/Miscellaneous/actin_quant_pipeline.py, lines 1–31 · score 0.87 · minor axis, scikit image, actin distribution, intensity profile, SciPy, NumPy
  3. [3] § Results › A subset of smORFs diverges in expression from the main ORF of the encoding gene ↔ Results/microproteins_dashboard.py, lines 395–439 · score 0.84 · altORF, uORF, psORF, lncRNA, iORF, dORF
  4. [4] § Methods › Transcriptional regulatory and RBP motif analyses ↔ Results/generate_supplemental_tables.py, lines 316–375 · score 0.82 · RBP motif, splice junction, gene symbol, Transcription factor, Benjamini Hochberg, Promoter
  5. [5] § Methods › Normalization, covariate adjustment and differential expression analysis of TMT proteomics data ↔ Code/Peptide_TMT_analysis/fragpipe_results_processing_scripts/TMT_regressed_corrected_matrix.R, lines 20–132 · score 0.80 · median coefficients, linear modeling, regressing, Bootstrap, age, matrix
  6. [6] § Results › A subset of smORFs diverges in expression from the main ORF of the encoding gene ↔ Results/generate_supplemental_tables.py, lines 256–315 · score 0.79 · additive model, Pearson correlation, uORF, Benjamini Hochberg, fold change, lncRNA
  7. [7] § Methods › Ribosome profiling re-analysis using Rp3 to assess translation of MS-identified MPs ↔ Results/microproteins_dashboard.py, lines 288–377 · score 0.79 · RiboCode, multi mapping, MS detected, translational evidence, ambiguous, Ribo seq
  8. [8] § Methods › Bulk RNA-seq profiling and differential expression analysis ↔ Code/Shortread_RNA_analysis/Shortread_deseq_processing_scripts/Main_smORF_LRT_analysis.R, lines 1–42 · score 0.76 · Pearson correlation, DESeq2, slope, psych, nested, additive
  9. [9] § Results › Unique MS evidence of MPs in aged human brains ↔ Results/generate_supplemental_tables.py, lines 196–255 · score 0.74 · razor peptides, confidence interval, unique peptides, Protein length, MS evidence, ShortStop
  10. [10] § Methods › DIA MS of human DLPFC tissue ↔ Code/gold_standard_filtering_criteria.py, lines 18–130 · score 0.74 · global pg, RiboCode, ShortStop, proteotypic, DIA, DDA
  11. [11] § Methods › Human postmortem brain cohorts ↔ Code/Shortread_RNA_analysis/Shortread_deseq_processing_scripts/RNA_differential_expression.R, lines 1–84 · score 0.73 · bulk RNA seq, asym AD, ROSMAP DLPFC, Clinical, age, variables
  12. [12] § Methods › Bulk RNA-seq profiling and differential expression analysis ↔ Code/Shortread_RNA_analysis/Shortread_deseq_processing_scripts/RNA_differential_expression.R, lines 1–84 · score 0.72 · DESeq2, limma, RNA seq, hg19, VST, matrices
  13. [13] § Results › Unique MS evidence of MPs in aged human brains ↔ Results/microproteins_dashboard.py, lines 288–377 · score 0.72 · ShortStop, MS detected, UniProt, Protein length, Ribo seq, RPKM
  14. [14] § Methods › Human postmortem brain cohorts ↔ Results/generate_supplemental_tables.py, lines 376–416 · score 0.71 · ROSMAP bulk RNA, DIA proteomics, asym AD, female, RNA seq, UCSD
  15. [15] § Methods › DIA MS of human DLPFC tissue ↔ Results/microproteins_dashboard.py, lines 1570–1661 · score 0.68 · global pg, RiboCode, proteotypic, DIA, DDA, Human
  16. [16] § Results › Unique MS evidence of MPs in aged human brains ↔ Results/microproteins_dashboard.py, lines 395–439 · score 0.66 · internal ORFs, lncRNA, short isoforms, TrEMBL, upstream, downstream
  17. [17] § Results › Unique MS evidence of MPs in aged human brains ↔ Results/microproteins_dashboard.py, lines 2293–2371 · score 0.65 · Swiss Prot matched, UniProt, TrEMBL, curation, smORFs, aa
  18. [18] § Methods › TMT-based proteomics and protein identification analysis ↔ Code/Peptide_TMT_analysis/fragpipe/fragpipe_round1.sh, lines 1–44 · score 0.64 · rescoring workflow, FragPipe, plex, variable, proteogenomic, batch
  19. [19] § Methods › Statistics and reproducibility ↔ Code/Peptide_TMT_analysis/fragpipe_results_processing_scripts/TMT_ANOVA.R, lines 86–165 · score 0.61 · post hoc, Benjamini Hochberg, broom, diagnoses, TMT
  20. [20] § Methods › Long-read transcriptome assembly and putative ORF database construction ↔ Results/generate_supplemental_tables.py, lines 316–375 · score 0.61 · Nanopore long, GTFtoFASTA, External, TrEMBL, RNA seq, Swiss Prot
  21. [21] § Results › Integrating Ribo-seq, MS and PROSIT provides an evidence framework for confidently identifying brain-expressed MPs ↔ Results/generate_supplemental_tables.py, lines 196–255 · score 0.61 · spectral angle, confidence tiers, protein identification, SA, coverage, grading
  22. [22] § Methods › smORF classification according to genomic context and isoform homology ↔ Results/microproteins_dashboard.py, lines 3916–3973 · score 0.61 · BLASTP alignment, bit scores, classified, smORFs, canonical, models
  23. [23] § Methods › Normalization, covariate adjustment and differential expression analysis of TMT proteomics data ↔ Code/Peptide_TMT_analysis/fragpipe_results_processing_scripts/TAMPOR_combined_rounds.R, lines 55–97 · score 0.60 · Median Polish, GIS, reporter, channel, rounds, batches
  24. [24] § Results › Integrating Ribo-seq, MS and PROSIT provides an evidence framework for confidently identifying brain-expressed MPs ↔ Code/Peptide_TMT_analysis/prosit/prosit_pipeline.py, lines 726–867 · score 0.60 · fragment ion, spectral angle, insufficient, SA, moderate, weak
  25. [25] § Results › Unique MS evidence of MPs in aged human brains ↔ Code/gold_standard_filtering_criteria.py, lines 18–130 · score 0.59 · MS detection, ShortStop, Ribo seq, MS evidence, Swiss Prot, DIA
  26. [26] § Results › Ribosome footprints and MS coverage provide orthogonal evidence for brain-expressed pseudogenes ↔ Code/Miscellaneous/actin_quant_pipeline.py, lines 1–31 · score 0.58 · actin intensity profile, actin distribution, fluorescence, Edge, metric, nuclear
  27. [27] § Methods › Statistics and reproducibility ↔ Code/Shortread_RNA_analysis/Shortread_deseq_processing_scripts/RNA_differential_expression.R, lines 119–176 · score 0.55 · DESeq2, blinded, Braak, CERAD, dplyr, Padj
  28. [28] § Results › Posttranscriptional splicing motif enrichment for genes encoding MPs ↔ Results/generate_supplemental_tables.py, lines 1–30 · score 0.53 · RBP motif enrichment, splice site, junctions, smORF, encoding
  29. [29] § Methods › Differential expression analysis of long-read RNA-seq data ↔ Code/Longread_RNA_analysis/ESPRESSO_data_processing_scripts/deseq_brain_espresso.r, lines 1–48 · score 0.50 · DESeq2, fold change, integers, sex, ESPRESSO, RNA

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 4,247 lines · 197 KB · no license · 7 matches

  1. import streamlit as st
  2. import pandas as pd
  3. import plotly.express as px
  4. import plotly.graph_objects as go
  5. from plotly.subplots import make_subplots
  6. from pathlib import Path
  7. from urllib.parse import quote, unquote
  8. import numpy as np
  9. import os
  10. import ast
  11. import re
  12. import html
  13. import hashlib
  14. import struct
  15. # scRNA Enrichment view is always enabled.
  16. INCLUDE_SCRNA = True
  17. # =============================================================================
  18. # REMOTE ASSET HOSTING (Hugging Face Datasets fallback)
  19. # =============================================================================
  20. # When the large figure directories (mirror_plots/, expression_profiles/,
  21. # smorf_cartoon_figures/) are not present locally (e.g. on Streamlit Cloud),
  22. # the dashboard streams the images from the public Hugging Face dataset.
  23. # Override via .streamlit/secrets.toml: assets_base_url = "https://..."
  24. DEFAULT_ASSETS_BASE_URL = "https://huggingface.co/datasets/brmiller/brain-microprotein-atlas/resolve/main"
  25. HF_REPO_ID = "brmiller/brain-microprotein-atlas"
  26. def _get_assets_base_url():
  27. try:
  28. return st.secrets.get("assets_base_url", DEFAULT_ASSETS_BASE_URL).rstrip("/")
  29. except Exception:
  30. return DEFAULT_ASSETS_BASE_URL
  31. @st.cache_data(show_spinner=False)
  32. def _list_remote_files(_repo=HF_REPO_ID):
  33. """Return the full file listing of the remote HF dataset (cached)."""
  34. try:
  35. from huggingface_hub import HfApi
  36. return HfApi().list_repo_files(repo_id=_repo, repo_type="dataset")
  37. except Exception as e:
  38. st.warning(f"Could not list remote assets ({e}); figures may be missing.")
  39. return []
  40. def _remote_url(rel_path):
  41. """Build a public URL for a path within the HF dataset."""
  42. return f"{_get_assets_base_url()}/{quote(rel_path)}"
  43. @st.cache_data(show_spinner=False)
  44. def _localize_hf_asset(rel_path, _repo=HF_REPO_ID):
  45. """Download a dataset asset and return a cached local path (or None).
  46. The HF dataset is stored on the Xet backend, so the plain
  47. resolve/main/<file> URLs 403 on a direct browser GET (what st.image would
  48. issue). hf_hub_download reconstructs the file via the xet client
  49. (anonymous, requires the hf_xet package) and caches it on disk, so we serve
  50. that local copy to st.image instead of the un-fetchable URL.
  51. """
  52. try:
  53. from huggingface_hub import hf_hub_download
  54. return hf_hub_download(repo_id=_repo, repo_type="dataset", filename=rel_path)
  55. except Exception:
  56. return None
  57. def _resolve_image_src(src):
  58. """Resolve an image/PDF source for st.image / iframe rendering.
  59. Local paths pass through unchanged. HF dataset URLs are Xet-backed and 403
  60. on a direct browser fetch, so download them server-side (cached) and return
  61. the local path; fall back to the original URL if the download fails.
  62. """
  63. if not src or not isinstance(src, str):
  64. return src
  65. if not src.startswith(("http://", "https://")):
  66. return src
  67. marker = "/resolve/main/"
  68. if "huggingface.co" in src and marker in src:
  69. rel = unquote(src.split(marker, 1)[1])
  70. local = _localize_hf_asset(rel)
  71. if local:
  72. return local
  73. return src
  74. # Reference render width for the entry page's full-width figures. The body
  75. # column is ~1130px on a 1600px window; images carry max-width:100% (see the
  76. # stylesheet) so they shrink on narrower windows rather than overflowing.
  77. ENTRY_FIGURE_WIDTH = 1130
  78. @st.cache_data(show_spinner=False)
  79. def _png_size(path):
  80. """(width, height) of a PNG, read from its 24-byte header. None if not a PNG.
  81. Cheap enough to call per render — no image library, no full decode.
  82. """
  83. try:
  84. with open(path, "rb") as fh:
  85. head = fh.read(24)
  86. except OSError:
  87. return None
  88. if len(head) < 24 or head[:8] != b"\x89PNG\r\n\x1a\n":
  89. return None
  90. w, h = struct.unpack(">II", head[16:24])
  91. return (int(w), int(h)) if w and h else None
  92. _MEDIABOX_RE = re.compile(
  93. rb"/MediaBox\s*\[\s*([\d.+-]+)\s+([\d.+-]+)\s+([\d.+-]+)\s+([\d.+-]+)\s*\]"
  94. )
  95. def _mediabox(blob):
  96. """(width, height) from the first /MediaBox in a byte blob, or None."""
  97. m = _MEDIABOX_RE.search(blob)
  98. if not m:
  99. return None
  100. try:
  101. x0, y0, x1, y1 = (float(v) for v in m.groups())
  102. except ValueError:
  103. return None
  104. w, h = abs(x1 - x0), abs(y1 - y0)
  105. return (w, h) if w and h else None
  106. @st.cache_data(show_spinner=False)
  107. def _pdf_size(path):
  108. """(width, height) in points of a PDF's first page, from its /MediaBox.
  109. The PNG path gets its aspect ratio from a 24-byte header; PDFs need the
  110. same thing so their figures can reserve height too. These are matplotlib
  111. exports that use compressed object streams, so /MediaBox is usually *not*
  112. findable in the raw bytes — inflating each Flate stream and searching the
  113. result is what actually finds it. No PDF library, no full parse; the files
  114. are ~60 KB, so brute force is cheap and the result is cached per path.
  115. """
  116. try:
  117. with open(path, "rb") as fh:
  118. data = fh.read()
  119. except OSError:
  120. return None
  121. if not data.startswith(b"%PDF"):
  122. return None
  123. # Uncompressed case first — cheapest, and covers PDFs from other tools.
  124. size = _mediabox(data)
  125. if size:
  126. return size
  127. import zlib
  128. for m in re.finditer(rb"stream\r?\n", data):
  129. start = m.end()
  130. end = data.find(b"endstream", start)
  131. if end < 0:
  132. continue
  133. try:
  134. inflated = zlib.decompress(data[start:end])
  135. except zlib.error:
  136. continue # not Flate, or not a self-contained stream — skip it
  137. size = _mediabox(inflated)
  138. if size:
  139. return size
  140. return None
  141. def _pdf_reserved(src, width=ENTRY_FIGURE_WIDTH, fallback_height=420):
  142. """Inline PDF in a height-reserved container, mirroring _image_reserved.
  143. An <iframe> does reserve its own box, but only at whatever height the call
  144. site guessed — so a figure toggled between a PNG and a PDF rendition jumped
  145. as the box resized. Deriving the height from the page's own aspect ratio
  146. keeps both renditions the same shape and keeps the anchor targets still.
  147. """
  148. resolved = _resolve_image_src(src)
  149. size = _pdf_size(resolved) if isinstance(resolved, str) else None
  150. height = int(round(width * size[1] / size[0])) if size else fallback_height
  151. with st.container(height=height + 8, border=False):
  152. _display_pdf_inline(resolved, height=height)
  153. def _image_reserved(src, width=ENTRY_FIGURE_WIDTH):
  154. """st.image, with the image's height reserved *before* it loads.
  155. An <img> with no intrinsic size known to the browser occupies zero height
  156. until it decodes, so everything below it jumps down when it arrives. On the
  157. Hugging Face Space the figures stream from the dataset rather than sitting on
  158. local disk, so that arrival is seconds late — long enough that clicking a
  159. section in the entry nav scrolls to a target that then slides out from under
  160. you (measured: 584px of drift; the browser will not compensate, scroll
  161. anchoring is off app-wide and re-enabling it made no difference).
  162. Reading the PNG header gives the aspect ratio server-side, so the height can
  163. be reserved with a fixed-size container and the page stops moving.
  164. """
  165. resolved = _resolve_image_src(src)
  166. size = _png_size(resolved) if isinstance(resolved, str) else None
  167. if not size:
  168. # Not a local PNG (e.g. an un-fetchable remote URL) — nothing to reserve.
  169. st.image(resolved, width=width)
  170. return
  171. # +8px so rounding can never leave the image taller than its own box, which
  172. # would give the container an inner scrollbar.
  173. height = int(round(width * size[1] / size[0])) + 8
  174. with st.container(height=height, border=False):
  175. st.image(resolved, width=width)
  176. # =============================================================================
  177. # PROSIT MIRROR PLOT CONFIGURATION
  178. # =============================================================================
  179. MIRROR_PLOT_BASE = Path(__file__).parent / "mirror_plots"
  180. QUALITY_LEVELS = ['Strong', 'Moderate', 'Weak', 'Insufficient', 'No PROSIT']
  181. QUALITY_EMOJI = {
  182. 'Strong': 'Strong',
  183. 'Moderate': 'Moderate',
  184. 'Weak': 'Weak',
  185. 'Insufficient': 'Insufficient',
  186. 'No PROSIT': 'No PROSIT',
  187. }
  188. # ── TMT significance tiers (best/most-stringent tier a microprotein reaches) ──
  189. # Tier 1 = strongest signal (Green) … Tier 4 = weakest (Red).
  190. # Tier 1: TMT q < 0.05 in ≥50% samples/condition
  191. # Tier 2: TMT q < 0.2 in ≥50% samples/condition
  192. # Tier 3: TMT q < 0.05 in ≥1 sample/condition
  193. # Tier 4: TMT q < 0.2 in ≥1 sample/condition
  194. TMT_TIER_EMOJI = {
  195. 'Tier 1': 'Tier 1',
  196. 'Tier 2': 'Tier 2',
  197. 'Tier 3': 'Tier 3',
  198. 'Tier 4': 'Tier 4',
  199. }
  200. # ── ROSMAP RNA-seq significance (Yellow = Exploratory, Green = Significant) ──
  201. RNA_SIG_EMOJI = {
  202. 'Significant': 'Significant',
  203. 'Exploratory': 'Exploratory',
  204. }
  205. # ── smORF decay class (NMD/NSD status × main-ORF disruption) ──────────────────
  206. # From the external NMD analysis shipped as
  207. # Code/data/smorf_trx_priority_color_class_assignments.tsv, whose `category`
  208. # column encodes the NMD × main-ORF-disruption 2×2 plus two non-stop-decay
  209. # (NSD) classes, as genome-browser track colors. NSD is decided per smORF over
  210. # the whole transcript set — one confident no-stop transcript and none with a
  211. # real stop — and overrides the 2×2, so it is a third decay state, not a fifth
  212. # color. Dots below match those track colors so the dashboard reads the same
  213. # as the BED in a browser.
  214. # Covers Salk/TrEMBL smORFs only — Swiss-Prot microproteins were never assessed
  215. # upstream, so they legitimately carry 'NA'.
  216. DECAY_CLASS_TSV = (Path(__file__).resolve().parent.parent / "Code" / "data" /
  217. "smorf_trx_priority_color_class_assignments.tsv")
  218. DECAY_CLASS_LEVELS = ['No NMD · mORF Shift', 'No NMD · mORF Intact',
  219. 'NMD · mORF Shift', 'NMD · mORF Intact',
  220. 'NSD · mORF Shift', 'NSD · mORF Intact', 'NA']
  221. DECAY_CLASS_FROM_CATEGORY = {
  222. 'dark_green': 'No NMD · mORF Shift',
  223. 'light_green': 'No NMD · mORF Intact',
  224. 'gold': 'NMD · mORF Shift',
  225. 'firebrick': 'NMD · mORF Intact',
  226. 'nsd_disrupted': 'NSD · mORF Shift',
  227. 'nsd_intact': 'NSD · mORF Intact',
  228. }
  229. DECAY_CLASS_EMOJI = {
  230. 'No NMD · mORF Shift': 'No NMD · mORF Shift',
  231. 'No NMD · mORF Intact': 'No NMD · mORF Intact',
  232. 'NMD · mORF Shift': 'NMD · mORF Shift',
  233. 'NMD · mORF Intact': 'NMD · mORF Intact',
  234. 'NSD · mORF Shift': 'NSD · mORF Shift',
  235. 'NSD · mORF Intact': 'NSD · mORF Intact',
  236. 'NA': 'NA',
  237. }
  238. DECAY_CLASS_HELP = (
  239. 'Predicted decay fate of the host transcript × whether the smORF shifts '
  240. 'the main ORF. NMD = stop trips the 50-nt rule; NSD = non-stop decay, an '
  241. 'ORF with no in-frame stop. NSD is judged over the whole transcript set — '
  242. 'it needs a confident no-stop transcript and no transcript carrying a real '
  243. 'stop — so it is scored ahead of, and overrides, the NMD × disruption 2×2. '
  244. 'NA = not assessed (Swiss-Prot microproteins were not covered by this '
  245. 'analysis).'
  246. )
  247. # =============================================================================
  248. # MAIN-TABLE COLUMN DESCRIPTIONS
  249. # =============================================================================
  250. # One line per display column name (as used in display_df / _column_groups).
  251. # Single source of truth for the "Explain the column/variables to me" expander —
  252. # shows only the descriptions for whatever columns are visible in the active
  253. # column-view tab.
  254. COLUMN_DESCRIPTIONS = {
  255. 'DB': 'Database source: 🔵 reviewed Swiss-Prot vs 🟠 unreviewed Salk/TrEMBL discovery.',
  256. 'sequence': 'The annotated microprotein amino-acid sequence, as originally called (from its annotated start through the stop codon).',
  257. 'NonATG Alternative Sequence(s)': "Predicted sequence(s) from alternative (non-AUG) initiation candidates (nonATG_predicted_sequence), raw list format; blank for ATG-start microproteins or non-ATG starts with no compatible alternative site.",
  258. 'NonATG Candidates': 'Alternative initiation candidates found by scanning, as codon@position (e.g. "CTG@23"); blank if the start is ATG or no alternative site was found.',
  259. 'UCSC': "Link to view this microprotein's genomic locus in the UCSC Genome Browser.",
  260. 'Parent Gene': 'Host gene symbol/name the smORF resides within or near.',
  261. 'General smORF Type': 'Broad smORF category (e.g. Upstream, Downstream, Short-Isoform, lncRNA-encoded, Swiss-Prot).',
  262. 'smORF Subtype': 'Specific smORF subtype/classification within its general type.',
  263. 'ORF Rules': 'Predicted decay fate of the host transcript × whether the smORF shifts the main ORF (no NMD/mORF shift, no NMD/mORF intact, NMD/mORF shift, NMD/mORF intact, NSD/mORF shift, NSD/mORF intact). NMD = stop trips the 50-nt rule; NSD = non-stop decay, an ORF with no in-frame stop — scored per smORF over the whole transcript set and overriding the NMD call. NA = not assessed; Swiss-Prot microproteins were not covered by this analysis.',
  264. 'Protein Length': 'Length of the microprotein in amino acids.',
  265. 'Start Codon': 'Coarse call: ATG (canonical) vs nonATG (annotated with a near-cognate/non-standard start).',
  266. 'Kozak Strength': 'Full Kozak consensus context strength around the start codon (weak/adequate/strong); Salk smORFs only.',
  267. 'Kozak Downstream Strength': 'Kozak context strength considering only the +4 downstream position.',
  268. 'Kozak Class': 'Canonical combined Kozak strength classification.',
  269. 'Kozak Window': 'Local sequence context around the start codon (lowercase = flanking, uppercase = start + downstream).',
  270. 'PhyloCSF Score': 'PhyloCSF evolutionary conservation score.',
  271. 'ShortStop': 'ShortStop ribosome-profiling classification label (e.g. SAM-Secreted, SAM-Intracellular).',
  272. 'ShortStop Score': 'ShortStop ML confidence score (0–1).',
  273. 'Annotation Method': 'How the microprotein was annotated: MS (mass-spec detected) or RiboCode-ShortStop (translation evidence without MS detection).',
  274. 'Spectra Quality': 'Best PROSIT spectral-match confidence tier (from the master Confidence column).',
  275. 'Nt-Acetylated': 'At least one tryptic peptide carries an N-terminal acetyl mark — direct evidence that this is a genuine protein N-terminus.',
  276. 'Nt-Acetyl Peptides': 'Tryptic peptide sequence(s) observed with an N-terminal acetyl mark.',
  277. 'Nt-Acetyl PSMs': 'Total Nt-acetylated PSMs summed across the Nt-acetylated peptides.',
  278. 'Nt-Acetyl PSM Fraction': 'Highest per-peptide fraction of that peptide’s PSMs carrying the Nt-acetyl mark (1.0 = every PSM acetylated).',
  279. 'TMT Tier': 'TMT significance tier: Tier 1 (strongest) … Tier 4.',
  280. 'RNA Significance': 'ROSMAP RNA-seq significance: Significant, Exploratory.',
  281. 'UniProt Annotation': 'UniProt curation annotation score (1–5★) for reviewed/TrEMBL entries; blank for novel microproteins not in UniProt.',
  282. 'NonATG Codon': 'The originally annotated (non-ATG) start codon.',
  283. 'NonATG Valid Initiator': 'Whether the annotated codon is itself a plausible translation initiator.',
  284. 'NonATG Has Supported Site': 'Whether current non-canonical initiation evidence/models support this annotated start actually being used (holistic verdict, distinct from Valid Initiator); Salk smORFs only.',
  285. 'NonATG Context Strength': 'Kozak context strength around the annotated non-ATG codon.',
  286. 'NonATG Alt Sites Found': 'Number of peptide-compatible alternative initiation sites found by scanning downstream of the annotated start.',
  287. 'Optimal Codon Tier': 'Strongest evidence tier among the alternative initiation candidates (Well-Established Near-Cognate is best).',
  288. 'Unique Spectral Counts (DDA)': 'Number of unique mass-spectrometry spectral counts (DDA).',
  289. 'Razor Counts (DDA)': 'Total razor spectral counts (shared peptides assigned to this protein).',
  290. 'TMT log2FC (50%)': 'TMT log2 fold-change (AD vs Control) — 50% missing-value threshold.',
  291. 'TMT t-stat (50%)': 'TMT t-statistic (AD vs Control) — 50% missing-value threshold.',
  292. 'TMT df (50%)': 'TMT t-test degrees of freedom — 50% missing-value threshold.',
  293. 'TMT CI low (50%)': 'TMT 95% CI lower bound on log2 fold-change — 50% missing-value threshold.',
  294. 'TMT CI high (50%)': 'TMT 95% CI upper bound on log2 fold-change — 50% missing-value threshold.',
  295. "TMT Cohen's d (50%)": "TMT Cohen's d effect size (AD vs Control) — 50% missing-value threshold.",
  296. 'TMT p-val (50%)': 'TMT raw p-value — 50% missing-value threshold.',
  297. 'TMT q-val (50%)': 'TMT q-value (BH-adjusted) — 50% missing-value threshold.',
  298. 'TMT log2FC (0%)': 'TMT log2 fold-change (AD vs Control) — 0% missing-value threshold (stringent).',
  299. 'TMT t-stat (0%)': 'TMT t-statistic (AD vs Control) — 0% missing-value threshold (stringent).',
  300. 'TMT df (0%)': 'TMT t-test degrees of freedom — 0% missing-value threshold (stringent).',
  301. 'TMT CI low (0%)': 'TMT 95% CI lower bound on log2 fold-change — 0% missing-value threshold (stringent).',
  302. 'TMT CI high (0%)': 'TMT 95% CI upper bound on log2 fold-change — 0% missing-value threshold (stringent).',
  303. "TMT Cohen's d (0%)": "TMT Cohen's d effect size (AD vs Control) — 0% missing-value threshold (stringent).",
  304. 'TMT p-val (0%)': 'TMT raw p-value — 0% missing-value threshold (stringent).',
  305. 'TMT q-val (0%)': 'TMT q-value (BH-adjusted) — 0% missing-value threshold (stringent).',
  306. 'MS Detect Control': 'MS detection rate in Control donors.',
  307. 'MS Detect AD': 'MS detection rate in AD donors.',
  308. 'Tryptic Peptides': 'Observed tryptic peptide sequence(s) identified by mass spec.',
  309. 'Tryptic Protein ID': 'Protein ID associated with the tryptic peptides.',
  310. 'Tryptic Start Positions': 'Start position(s) of the tryptic peptide(s) within the annotated sequence.',
  311. 'Tryptic End Positions': 'End position(s) of the tryptic peptide(s) within the annotated sequence.',
  312. 'ROSMAP log2FC': 'ROSMAP short-read RNA-seq log2 fold-change (AD vs Control).',
  313. 'ROSMAP padj': 'ROSMAP short-read RNA-seq adjusted p-value.',
  314. 'MSBB log2FC': 'MSBB short-read RNA-seq log2 fold-change (AD vs Control).',
  315. 'MSBB padj': 'MSBB short-read RNA-seq adjusted p-value.',
  316. 'ROSMAP Corr NonAD': 'ROSMAP: correlation between Main ORF and smORF transcripts (non-AD donors).',
  317. 'ROSMAP Corr AD': 'ROSMAP: correlation between Main ORF and smORF transcripts (AD donors).',
  318. 'MSBB Corr NonAD': 'MSBB: correlation between Main ORF and smORF transcripts (non-AD donors).',
  319. 'MSBB Corr AD': 'MSBB: correlation between Main ORF and smORF transcripts (AD donors).',
  320. 'RNA_LRT_Add_P': 'ROSMAP RNA-seq LRT additive-model p-value.',
  321. 'RNA_LRT_Int_P': 'ROSMAP RNA-seq LRT interaction-model p-value.',
  322. 'RP3 Default': 'RP3 default pipeline ribosome-profiling result.',
  323. 'RP3 MM+Amb': 'RP3 multi-mapping + ambiguous-reads result.',
  324. 'RP3 Amb': 'RP3 ambiguous-reads result.',
  325. 'RP3 MM': 'RP3 multi-mapping result.',
  326. 'RiboCode': 'RiboCode ORF-detection result.',
  327. 'P-site % Frame 0': 'Percent of the ORF\'s P-sites at codon position 1 (in frame). Unreviewed smORFs matched by sequence, TrEMBL by CDS coordinates, Swiss-Prot by gene name + length.',
  328. 'P-site % Frame 1': 'Percent of the ORF\'s P-sites at codon position 2.',
  329. 'P-site % Frame 2': 'Percent of the ORF\'s P-sites at codon position 3.',
  330. 'P-site Frame 0 RPKM': 'Frame-0 P-sites per kb of CDS per million P-sites (41 adult Ribo-seq libraries).',
  331. 'BLAST UniProt Match': 'UniProt accession of the best BLASTp hit.',
  332. 'BLAST % Match': 'BLASTp percent identity to the best UniProt hit.',
  333. 'BLAST Aln Length': 'BLASTp alignment length (residues).',
  334. 'BLAST E-value': 'BLASTp expectation value of the best UniProt hit.',
  335. 'BLAST Bit Score': 'BLASTp bit score of the best UniProt hit.',
  336. }
  337. # ── UniProt annotation score (1–5 curation completeness, rendered as gold stars) ──
  338. def _annotation_stars(score):
  339. """Render a 1–5 UniProt annotation score as gold stars ('★★★☆☆'); '—' if missing."""
  340. if score is None or pd.isna(score):
  341. return '\u2014'
  342. try:
  343. n = int(round(float(score)))
  344. except (ValueError, TypeError):
  345. return '\u2014'
  346. n = max(0, min(5, n))
  347. return '\u2605' * n + '\u2606' * (5 - n)
  348. # =============================================================================
  349. # smORF CARTOON FIGURE CONFIGURATION
  350. # =============================================================================
  351. SMORF_CARTOON_DIR = Path(__file__).parent / "smorf_cartoon_figures"
  352. SEQ_TO_COORDS_FILE = Path(__file__).parent.parent / "GTF_and_BED_files" / "Unreviewed_Brain_Microproteins_mapping_coordinates_to_sequences.tsv"
  353. # =============================================================================
  354. # EXPRESSION PROFILE CONFIGURATION
  355. # =============================================================================
  356. EXPRESSION_PROFILE_DIR = Path(__file__).parent / "expression_profiles"
  357. GENOME_FILES_DIR = Path(__file__).parent.parent / "GTF_and_BED_files"
  358. # =============================================================================
  359. # smORF TYPE GROUPINGS
  360. # =============================================================================
  361. # smORF-type -> parent group. Kept in sync with the figure generators'
  362. # map_smorf_label (revision_figures/_stats_figure{1..4}): udORF and isoORF are
  363. # grouped as TrEMBL/altORF and Short-Isoform respectively.
  364. SMORF_PARENT_GROUPS = {
  365. 'Upstream': ['uORF', 'uoORF', 'uaORF', 'uaoORF'],
  366. 'Downstream': ['dORF', 'doORF', 'daORF', 'daoORF'],
  367. 'Internal ORF': ['iORF', 'oORF'],
  368. 'TrEMBL/AltORF': ['eORF', 'udORF', 'TrEMBL'],
  369. 'Short-Isoform': ['Iso', 'D-Iso', 'N-Iso', 'isoORF'],
  370. 'lncRNA': ['lncRNA'],
  371. 'psORF': ['psORF'],
  372. 'Swiss-Prot': ['Swiss-Prot-MP'],
  373. }
  374. SMORF_CHILD_TO_PARENT = {
  375. child: parent
  376. for parent, children in SMORF_PARENT_GROUPS.items()
  377. for child in children
  378. }
  379. SMORF_DISPLAY_LABEL = {}
  380. # =============================================================================
  381. # scRNA CELL-TYPE ENRICHMENT CONFIGURATION
  382. # =============================================================================
  383. # Per-microprotein cell-type panel in the ID card, the single-gene analogue of
  384. # the cell-type heatmaps in Code/scRNAseq_summary_merging_analysis/
  385. # scRNAseq_summary.R: same diverging log2FC fill and the same
  386. # cell_type_general grouping, but one microprotein per plot instead of an
  387. # aggregate over smORF types.
  388. SCRNA_SUMMARY_FILE = Path(__file__).parent / "scRNA_Enrichment" / "scRNA_Enrichment_summary.csv"
  389. # Every tested (microprotein, cell type) pair, including the non-significant
  390. # ones the summary above drops. Written by the same R script; the panel needs it
  391. # to distinguish "tested, not significant" from "not tested".
  392. SCRNA_ALL_CELLTYPES_FILE = Path(__file__).parent / "scRNA_Enrichment" / "scRNA_Enrichment_all_celltypes.csv.gz"
  393. # Facet order from the R script's `ordered_cell_types`, with the catch-all last.
  394. SCRNA_GENERAL_ORDER = ['Astr', 'Exc Neur', 'Immune', 'Inh Neur', 'Oli', 'Vas', 'Other']
  395. # Okabe-Ito (colorblind-safe) — one hue per general cell type. Carries the
  396. # grouping that facet_grid(rows = vars(cell_type_general)) carries in the R
  397. # figure, which a single-column plotly heatmap has no strips for.
  398. SCRNA_GENERAL_COLORS = {
  399. 'Astr': '#0072B2',
  400. 'Exc Neur': '#D55E00',
  401. 'Immune': '#009E73',
  402. 'Inh Neur': '#CC79A7',
  403. 'Oli': '#E69F00',
  404. 'Vas': '#56B4E9',
  405. 'Other': '#999999',
  406. }
  407. # scale_fill_gradient2(low = "#4575b4", mid = "white", high = "#d73027") from
  408. # the R heatmaps, as a plotly colorscale.
  409. SCRNA_FC_COLORSCALE = [[0.0, '#4575b4'], [0.5, '#ffffff'], [1.0, '#d73027']]
  410. # The R heatmaps squish mean log2FC to ±0.3. Single-gene effects run larger, so
  411. # the panel expands past this floor when the microprotein needs it (see
  412. # _scrna_fc_limit) rather than saturating every tile.
  413. SCRNA_FC_FLOOR = 0.3
  414. # =============================================================================
  415. # UCSC CUSTOM SESSION CONFIGURATION
  416. # =============================================================================
  417. CUSTOM_UCSC_SESSION = "AD Dark Microproteome"
  418. CUSTOM_UCSC_USERNAME = "brmiller"
  419. CUSTOM_UCSC_SESSION_URL = "https://genome.ucsc.edu/s/brmiller/AD%20Dark%20Microproteome"
  420. COLORS = {
  421. 'swiss_prot': '#74a2b7',
  422. 'unreviewed': '#ed8651',
  423. }
  424. # =============================================================================
  425. # HELPER FUNCTIONS
  426. # =============================================================================
  427. # ── Entry-page addressing ────────────────────────────────────────────────────
  428. # `sequence` is the de-duplication key of the merged master table (see the
  429. # drop_duplicates in load_and_merge_all_data), so it is the microprotein's
  430. # identity. Sequences are far too long for a URL, so the entry page is
  431. # addressed by a short, stable digest of the sequence: ?mp=<entry id>.
  432. ENTRY_QUERY_PARAM = "mp"
  433. def _entry_id(seq):
  434. """Stable short accession-like id for a microprotein sequence."""
  435. s = str(seq or '')
  436. if not s:
  437. return ''
  438. return hashlib.sha1(s.encode('utf-8')).hexdigest()[:10]
  439. @st.cache_data
  440. def build_entry_id_index(_df, cache_key):
  441. """Map entry id -> sequence, over the *unfiltered* set.
  442. Built from the post-load frame rather than the filtered view so that an
  443. entry URL resolves regardless of which filters happen to be active.
  444. `_df` is underscore-prefixed (excluded from the cache key, per the
  445. convention used by the other index builders) to skip re-hashing a large
  446. frame on every rerun; `cache_key` — the row count — is what actually
  447. invalidates the entry, so pass it explicitly.
  448. """
  449. if _df is None or 'sequence' not in _df.columns:
  450. return {}
  451. seqs = _df['sequence'].dropna().astype(str).unique()
  452. return {_entry_id(s): s for s in seqs}
  453. @st.cache_data
  454. def build_seq_to_coords_index(_path=str(SEQ_TO_COORDS_FILE)):
  455. """Build sequence-to-coordinates lookup from the mapping file."""
  456. p = Path(_path)
  457. if not p.exists():
  458. return {}
  459. df = pd.read_csv(p, sep='\t')
  460. return dict(zip(df['sequence'], df['genomic_coordinates']))
  461. @st.cache_data
  462. def build_mirror_plot_index(_base=str(MIRROR_PLOT_BASE)):
  463. """Build index of available mirror plot images by peptide.
  464. Local-first: scans the on-disk directory if present, otherwise pulls the
  465. file listing from the Hugging Face dataset and builds URL-based entries.
  466. """
  467. base = Path(_base)
  468. index = {}
  469. def _add_entry(peptide, quality, parts, filepath):
  470. charge = parts[1].replace("z", "") if len(parts) > 1 else "?"
  471. scan = parts[-1] if len(parts) > 2 else "?"
  472. if peptide not in index:
  473. index[peptide] = {'best_quality': quality, 'plots': []}
  474. index[peptide]['plots'].append({
  475. 'quality': quality,
  476. 'charge': charge,
  477. 'scan': scan,
  478. 'filepath': filepath,
  479. 'peptide': peptide,
  480. })
  481. if base.exists():
  482. for quality in QUALITY_LEVELS:
  483. quality_dir = base / quality
  484. if not quality_dir.exists():
  485. continue
  486. for img_file in quality_dir.glob("*.png"):
  487. parts = img_file.stem.split("_")
  488. if len(parts) >= 3:
  489. _add_entry(parts[0], quality, parts, str(img_file))
  490. return index
  491. # Remote fallback
  492. valid_qualities = set(QUALITY_LEVELS)
  493. for rel in _list_remote_files():
  494. if not rel.startswith("mirror_plots/") or not rel.endswith(".png"):
  495. continue
  496. segs = rel.split("/")
  497. if len(segs) != 3:
  498. continue
  499. _, quality, fname = segs
  500. if quality not in valid_qualities:
  501. continue
  502. parts = Path(fname).stem.split("_")
  503. if len(parts) >= 3:
  504. _add_entry(parts[0], quality, parts, _remote_url(rel))
  505. return index
  506. @st.cache_data
  507. def build_expression_profile_index(_base=str(EXPRESSION_PROFILE_DIR)):
  508. """Build index of expression profile images keyed by genomic coordinates.
  509. Local-first; falls back to HF dataset file listing if local dir missing.
  510. """
  511. base = Path(_base)
  512. index = {}
  513. def _add(stem, coupling, filepath):
  514. match = re.search(r'(chr\w+)_(\d+)-(\d+)$', stem)
  515. if not match:
  516. return
  517. chrom, start, end = match.group(1), match.group(2), match.group(3)
  518. coords_key = f"{chrom}:{start}-{end}"
  519. gene_name = stem[:stem.rfind(f'_{chrom}')]
  520. index.setdefault(coords_key, []).append({
  521. 'filepath': filepath,
  522. 'coupling': coupling,
  523. 'gene': gene_name,
  524. })
  525. if base.exists():
  526. for subdir in ['coupled', 'non_coupled']:
  527. subpath = base / subdir
  528. if not subpath.exists():
  529. continue
  530. seen_stems = set()
  531. for img_file in list(subpath.glob("*.png")) + list(subpath.glob("*.pdf")):
  532. if img_file.stem in seen_stems:
  533. continue
  534. seen_stems.add(img_file.stem)
  535. _add(img_file.stem, subdir, str(img_file))
  536. return index
  537. # Remote fallback: prefer PNG over PDF per stem
  538. by_stem = {} # stem -> (rel_path, coupling, ext_priority)
  539. for rel in _list_remote_files():
  540. if not rel.startswith("expression_profiles/"):
  541. continue
  542. segs = rel.split("/")
  543. if len(segs) != 3:
  544. continue
  545. _, subdir, fname = segs
  546. if subdir not in ('coupled', 'non_coupled'):
  547. continue
  548. stem, ext = Path(fname).stem, Path(fname).suffix.lower()
  549. if ext not in ('.png', '.pdf'):
  550. continue
  551. priority = 0 if ext == '.png' else 1
  552. existing = by_stem.get(stem)
  553. if existing is None or priority < existing[2]:
  554. by_stem[stem] = (rel, subdir, priority)
  555. for stem, (rel, subdir, _) in by_stem.items():
  556. _add(stem, subdir, _remote_url(rel))
  557. return index
  558. def _display_pdf_inline(filepath, height=500):
  559. """Display a PDF inline. Local files are base64-embedded; URLs are streamed."""
  560. filepath = _resolve_image_src(filepath)
  561. try:
  562. if isinstance(filepath, str) and filepath.startswith(("http://", "https://")):
  563. src = filepath
  564. else:
  565. import base64
  566. with open(filepath, "rb") as f:
  567. b64 = base64.b64encode(f.read()).decode()
  568. src = f"data:application/pdf;base64,{b64}"
  569. st.markdown(
  570. f'<iframe src="{src}" '
  571. f'width="100%" height="{height}px" style="border:none; border-radius:8px;"></iframe>',
  572. unsafe_allow_html=True,
  573. )
  574. except Exception as e:
  575. st.warning(f"Could not display PDF: {e}")
  576. def get_spectra_quality(tryptic_peptides_str, mirror_index):
  577. """Get best spectra quality for a microprotein's tryptic peptides."""
  578. if not tryptic_peptides_str or pd.isna(tryptic_peptides_str):
  579. return 'No PROSIT'
  580. if not mirror_index:
  581. return 'No PROSIT'
  582. try:
  583. peptides = ast.literal_eval(str(tryptic_peptides_str))
  584. if isinstance(peptides, str):
  585. peptides = [peptides]
  586. except (ValueError, SyntaxError):
  587. peptides = []
  588. # Default to 'No PROSIT': a microprotein with no PROSIT grade and no matching
  589. # mirror-plot spectrum has no spectral match, so it must NOT be counted as
  590. # 'Insufficient' (that tier is reserved for genuine low-quality PROSIT hits).
  591. best = 'No PROSIT'
  592. for pep in peptides:
  593. pep = pep.strip()
  594. if pep in mirror_index:
  595. q = mirror_index[pep]['best_quality']
  596. if QUALITY_LEVELS.index(q) < QUALITY_LEVELS.index(best):
  597. best = q
  598. return best
  599. def get_matching_mirror_plots(tryptic_peptides_str, mirror_index):
  600. """Get all mirror plots for a microprotein's tryptic peptides."""
  601. if not tryptic_peptides_str or pd.isna(tryptic_peptides_str) or not mirror_index:
  602. return []
  603. try:
  604. peptides = ast.literal_eval(str(tryptic_peptides_str))
  605. if isinstance(peptides, str):
  606. peptides = [peptides]
  607. except (ValueError, SyntaxError):
  608. return []
  609. matches = []
  610. for pep in peptides:
  611. pep = pep.strip()
  612. if pep in mirror_index:
  613. matches.extend(mirror_index[pep]['plots'])
  614. matches.sort(key=lambda x: QUALITY_LEVELS.index(x['quality']))
  615. return matches
  616. def create_ucsc_link(row, custom_session_id=None):
  617. """Create UCSC Genome Browser link."""
  618. def add_session_to_url(url, session_id):
  619. if session_id and 'genome.ucsc.edu' in url and CUSTOM_UCSC_SESSION_URL:
  620. position_match = re.search(r'position=([^&]+)', url)
  621. if position_match:
  622. position = position_match.group(1)
  623. sn = quote(CUSTOM_UCSC_SESSION)
  624. return (f"https://genome.ucsc.edu/cgi-bin/hgTracks?db=hg38"
  625. f"&hgS_doOtherUser=submit"
  626. f"&hgS_otherUserName={CUSTOM_UCSC_USERNAME}"
  627. f"&hgS_otherUserSessionName={sn}"
  628. f"&position={position}")
  629. else:
  630. sn = quote(CUSTOM_UCSC_SESSION)
  631. return (f"https://genome.ucsc.edu/cgi-bin/hgTracks?db=hg38"
  632. f"&hgS_doOtherUser=submit"
  633. f"&hgS_otherUserName={CUSTOM_UCSC_USERNAME}"
  634. f"&hgS_otherUserSessionName={sn}")
  635. return url
  636. if 'CLICK_UCSC' in row and pd.notna(row['CLICK_UCSC']):
  637. click_ucsc = str(row['CLICK_UCSC'])
  638. if click_ucsc.startswith('=HYPERLINK('):
  639. match = re.search(r'=HYPERLINK\("([^"]+)"', click_ucsc)
  640. if match:
  641. return add_session_to_url(match.group(1), custom_session_id)
  642. elif click_ucsc.startswith('http'):
  643. return add_session_to_url(click_ucsc, custom_session_id)
  644. coords = None
  645. if 'smORF Coordinates' in row and pd.notna(row['smORF Coordinates']):
  646. coords = row['smORF Coordinates']
  647. elif 'genomic_coordinates' in row and pd.notna(row['genomic_coordinates']):
  648. coords = row['genomic_coordinates']
  649. if coords and coords.startswith('chr'):
  650. try:
  651. chrom, pos = coords.split(':')
  652. start, end = pos.split('-')
  653. position = f"{chrom}:{start}-{end}"
  654. if custom_session_id and CUSTOM_UCSC_SESSION_URL:
  655. sn = quote(CUSTOM_UCSC_SESSION)
  656. return (f"https://genome.ucsc.edu/cgi-bin/hgTracks?db=hg38"
  657. f"&hgS_doOtherUser=submit"
  658. f"&hgS_otherUserName={CUSTOM_UCSC_USERNAME}"
  659. f"&hgS_otherUserSessionName={sn}"
  660. f"&position={position}")
  661. else:
  662. return f"https://genome.ucsc.edu/cgi-bin/hgTracks?db=hg38&position={position}"
  663. except Exception:
  664. pass
  665. return None
  666. # =============================================================================
  667. # PAGE CONFIG (must be first Streamlit command)
  668. # =============================================================================
  669. st.set_page_config(
  670. page_title="Brain Microprotein Dashboard",
  671. page_icon="\U0001f9ec",
  672. layout="wide",
  673. # Filters are visible on arrival so the sidebar is discoverable rather than
  674. # hidden behind the toggle. Streamlit applies this on first load only; once
  675. # a visitor opens or closes the sidebar, their choice sticks for the session.
  676. # (The entry page puts nothing in the sidebar, so none is rendered there.)
  677. initial_sidebar_state="expanded",
  678. )
  679. # =============================================================================
  680. # GLASSMORPHISM CSS
  681. # =============================================================================
  682. st.markdown(f"""
  683. <style>
  684. /* ── App gradient background ── */
  685. .stApp, .stApp > header, [data-testid="stHeader"] {{
  686. background: linear-gradient(135deg,
  687. #1a202c 0%, #2d3748 25%, #1a202c 50%, #2d3748 75%, #1a202c 100%) !important;
  688. }}
  689. /* ── Always reserve the vertical scrollbar gutter. Prevents the layout
  690. "shake"/jitter loop when tall content (e.g. PROSIT mirror plots) toggles
  691. the scrollbar on/off, which resizes width-stretched images and re-toggles
  692. it endlessly. ── */
  693. html, body {{
  694. overflow-y: scroll !important;
  695. }}
  696. /* Streamlit's real scroll container — keep the scrollbar gutter reserved and
  697. disable scroll anchoring so image loads don't nudge the viewport. */
  698. [data-testid="stAppViewContainer"], [data-testid="stMain"], section.main {{
  699. scrollbar-gutter: stable !important;
  700. overflow-anchor: none !important;
  701. }}
  702. /* NOTE: scroll anchoring is NOT the lever for the entry-page nav. Re-enabling it
  703. here (on the scroller and every ancestor) was measured and did not compensate
  704. for late-loading figures — 584px of drift either way. The entry page instead
  705. reserves each figure's height up front so the layout never shifts; see
  706. _image_reserved. */
  707. /* Cap any figure image so a fixed-width plot never forces a horizontal scrollbar. */
  708. [data-testid="stImage"] img {{
  709. max-width: 100% !important;
  710. height: auto !important;
  711. }}
  712. /* ── Main container glass ── */
  713. .main .block-container {{
  714. background: rgba(26, 32, 44, 0.82) !important;
  715. backdrop-filter: blur(12px) !important;
  716. border-radius: 16px !important;
  717. border: 1px solid rgba(116, 162, 183, 0.25) !important;
  718. box-shadow: 0 8px 32px rgba(0,0,0,0.35) !important;
  719. padding: 2rem !important;
  720. margin-top: 1rem !important;
  721. }}
  722. /* ── Sidebar gradient ── */
  723. section[data-testid="stSidebar"] > div:first-child {{
  724. background: linear-gradient(180deg,
  725. rgba(26,32,44,0.96) 0%, rgba(45,55,72,0.96) 50%, rgba(26,32,44,0.96) 100%) !important;
  726. backdrop-filter: blur(10px) !important;
  727. border-right: 1px solid rgba(116,162,183,0.35) !important;
  728. }}
  729. /* ── Glass card (reusable) ── */
  730. .glass-card {{
  731. background: rgba(255,255,255,0.06) !important;
  732. backdrop-filter: blur(14px) !important;
  733. border: 1px solid rgba(116,162,183,0.25) !important;
  734. border-radius: 14px !important;
  735. padding: 1.1rem 1.3rem !important;
  736. box-shadow: 0 4px 18px rgba(0,0,0,0.22) !important;
  737. margin-bottom: 0.6rem !important;
  738. }}
  739. .glass-card-section {{
  740. background: rgba(255,255,255,0.04) !important;
  741. backdrop-filter: blur(10px) !important;
  742. border: 1px solid rgba(116,162,183,0.18) !important;
  743. border-radius: 12px !important;
  744. padding: 1rem 1.2rem !important;
  745. margin-bottom: 0.75rem !important;
  746. }}
  747. /* ── ID card section header ── */
  748. .id-section-header {{
  749. font-size: 0.95rem;
  750. font-weight: 600;
  751. margin-bottom: 0.6rem;
  752. padding-bottom: 0.35rem;
  753. border-bottom: 1px solid rgba(116,162,183,0.22);
  754. color: rgba(255,255,255,0.92) !important;
  755. }}
  756. /* ── ID card field ── */
  757. .id-field-label {{
  758. font-size: 0.72rem;
  759. text-transform: uppercase;
  760. letter-spacing: 0.06em;
  761. color: rgba(255,255,255,0.50) !important;
  762. margin-bottom: 0.15rem;
  763. }}
  764. .id-field-value {{
  765. font-size: 1.05rem;
  766. font-weight: 500;
  767. color: #ffffff !important;
  768. margin-bottom: 0.7rem;
  769. }}
  770. /* Sequence-like values (coordinates, Kozak windows) — narrower and monospaced
  771. so they fit the grid cell without wrapping. */
  772. .id-field-value.id-field-mono {{
  773. font-family: monospace;
  774. font-size: 0.82rem;
  775. }}
  776. .id-field-value a {{
  777. color: {COLORS['swiss_prot']} !important;
  778. text-decoration: none !important;
  779. border-bottom: 1px dotted rgba(116,162,183,0.6);
  780. }}
  781. .id-field-value a:hover {{
  782. color: #ffffff !important;
  783. }}
  784. /* ── Command-center header ── */
  785. .cmd-header {{
  786. background: linear-gradient(135deg, {COLORS['swiss_prot']} 0%, {COLORS['unreviewed']} 100%);
  787. border-radius: 14px;
  788. padding: 1.2rem 1.6rem 1rem 1.6rem;
  789. margin-bottom: 1rem;
  790. box-shadow: 0 6px 24px rgba(0,0,0,0.25);
  791. position: relative;
  792. overflow: hidden;
  793. }}
  794. .cmd-header::before {{
  795. content: '';
  796. position: absolute;
  797. inset: 0;
  798. background: radial-gradient(ellipse at 20% 50%, rgba(255,255,255,0.12) 0%, transparent 70%);
  799. pointer-events: none;
  800. }}
  801. .cmd-top {{
  802. display: flex;
  803. justify-content: space-between;
  804. align-items: baseline;
  805. margin-bottom: 0.8rem;
  806. position: relative;
  807. }}
  808. .cmd-title {{
  809. font-size: 1.45rem;
  810. font-weight: 700;
  811. color: #ffffff !important;
  812. margin: 0;
  813. letter-spacing: -0.01em;
  814. }}
  815. .cmd-subtitle {{
  816. font-size: 0.78rem;
  817. color: rgba(255,255,255,0.65) !important;
  818. text-align: right;
  819. line-height: 1.4;
  820. }}
  821. .cmd-metrics {{
  822. display: grid;
  823. grid-template-columns: repeat(6, 1fr);
  824. gap: 0.55rem;
  825. position: relative;
  826. }}
  827. .cmd-stat {{
  828. background: rgba(0,0,0,0.18);
  829. backdrop-filter: blur(12px);
  830. border: 1px solid rgba(255,255,255,0.12);
  831. border-radius: 10px;
  832. padding: 0.6rem 0.5rem;
  833. text-align: center;
  834. transition: background 0.2s, transform 0.2s, box-shadow 0.2s;
  835. }}
  836. .cmd-stat:hover {{
  837. background: rgba(0,0,0,0.28);
  838. transform: translateY(-2px);
  839. box-shadow: 0 0 14px rgba(116,162,183,0.30), 0 4px 12px rgba(0,0,0,0.3);
  840. }}
  841. .cmd-stat .stat-label {{
  842. font-size: 0.65rem;
  843. text-transform: uppercase;
  844. letter-spacing: 0.06em;
  845. color: rgba(255,255,255,0.6);
  846. margin-bottom: 0.2rem;
  847. }}
  848. .cmd-stat .stat-val {{
  849. font-size: 1.45rem;
  850. font-weight: 700;
  851. color: #ffffff;
  852. line-height: 1.1;
  853. }}
  854. .cmd-stat .stat-sub {{
  855. font-size: 0.62rem;
  856. color: rgba(255,255,255,0.45);
  857. margin-top: 0.1rem;
  858. }}
  859. .cmd-stat.st-total .stat-val {{
  860. background: linear-gradient(135deg, #b8d8e8, #f5c4a1);
  861. -webkit-background-clip: text;
  862. -webkit-text-fill-color: transparent;
  863. background-clip: text;
  864. }}
  865. .cmd-stat.st-swiss {{
  866. border-color: rgba(116,162,183,0.45);
  867. }}
  868. .cmd-stat.st-swiss .stat-val {{
  869. color: #c5dce7;
  870. }}
  871. .cmd-stat.st-noncan {{
  872. border-color: rgba(237,134,81,0.45);
  873. }}
  874. .cmd-stat.st-noncan .stat-val {{
  875. color: #f5c4a1;
  876. }}
  877. .cmd-stat.st-swiss:hover {{
  878. box-shadow: 0 0 20px rgba(116,162,183,0.55), 0 4px 12px rgba(0,0,0,0.3) !important;
  879. }}
  880. .cmd-stat.st-noncan:hover {{
  881. box-shadow: 0 0 20px rgba(237,134,81,0.55), 0 4px 12px rgba(0,0,0,0.3) !important;
  882. }}
  883. /* ── Significance color helpers ── */
  884. .sig-up {{ color: #48bb78 !important; font-weight: 600; }}
  885. .sig-down {{ color: #fc8181 !important; font-weight: 600; }}
  886. .sig-ns {{ color: rgba(255,255,255,0.40) !important; }}
  887. .sig-yes {{ color: #68d391 !important; font-weight: 600; }}
  888. .sig-na {{ color: rgba(255,255,255,0.30) !important; }}
  889. /* ── Swiss-Prot / Unreviewed badges ── */
  890. .badge-swiss {{
  891. display: inline-block;
  892. background: {COLORS['swiss_prot']};
  893. color: #fff !important;
  894. padding: 2px 10px;
  895. border-radius: 12px;
  896. font-size: 0.78rem;
  897. font-weight: 600;
  898. }}
  899. .badge-unreviewed {{
  900. display: inline-block;
  901. background: {COLORS['unreviewed']};
  902. color: #fff !important;
  903. padding: 2px 10px;
  904. border-radius: 12px;
  905. font-size: 0.78rem;
  906. font-weight: 600;
  907. }}
  908. /* ═══ Entry page (UniProt-style detail view) ═══════════════════════════════ */
  909. /* Section picker. This used to be a list of <a href="#sec-…"> anchors, which
  910. worked locally but was dead on the Hugging Face Space: the huggingface.co
  911. wrapper embeds the app in a cross-origin iframe sized to the app's FULL
  912. content height, so the element the user actually scrolls is the *parent*
  913. huggingface.co document. A fragment link inside the iframe can only scroll
  914. scrollers inside that iframe, and same-origin policy forbids it touching the
  915. parent — measured on the live Space: parent scrollTop stayed 0 on every
  916. click, and the target landed at 443px instead of 72px.
  917. Bounding the app's own scroller does fix the landing (verified: exactly
  918. 72px), but HF then grew the iframe to 7509px — it sizes from a postMessage,
  919. not from the document, so CSS cannot settle it. Selecting a section and
  920. re-rendering needs no scrolling at all, so it behaves identically in the
  921. wrapper, on the direct *.hf.space URL, on Streamlit Cloud and locally.
  922. Stick the element container that *holds* the picker, not the column or its
  923. vertical block: those stretch to the full height of the row (i.e. of the
  924. entry body), and a sticky box as tall as its scroll parent has nowhere to
  925. travel, so it never appears to stick. */
  926. [data-testid="stElementContainer"]:has(.entry-nav-anchor) + [data-testid="stElementContainer"] {{
  927. position: sticky;
  928. top: 3.5rem;
  929. z-index: 5;
  930. background: rgba(255,255,255,0.04);
  931. backdrop-filter: blur(10px);
  932. border: 1px solid rgba(116,162,183,0.18);
  933. border-radius: 12px;
  934. padding: 0.6rem 0.5rem;
  935. }}
  936. /* Radio options styled as the old nav links. */
  937. [data-testid="stElementContainer"]:has(.entry-nav-anchor) + [data-testid="stElementContainer"]
  938. [role="radiogroup"] > label {{
  939. display: block;
  940. padding: 0.30rem 0.55rem;
  941. margin-bottom: 0.1rem;
  942. border-left: 2px solid transparent;
  943. border-radius: 0 6px 6px 0;
  944. transition: background 0.12s ease, border-color 0.12s ease;
  945. }}
  946. [data-testid="stElementContainer"]:has(.entry-nav-anchor) + [data-testid="stElementContainer"]
  947. [role="radiogroup"] > label:hover {{
  948. background: rgba(116,162,183,0.16);
  949. border-left-color: {COLORS['swiss_prot']};
  950. }}
  951. [data-testid="stElementContainer"]:has(.entry-nav-anchor) + [data-testid="stElementContainer"]
  952. [role="radiogroup"] > label:has(input:checked) {{
  953. background: rgba(116,162,183,0.20);
  954. border-left-color: {COLORS['unreviewed']};
  955. }}
  956. [data-testid="stElementContainer"]:has(.entry-nav-anchor) + [data-testid="stElementContainer"]
  957. [role="radiogroup"] > label p {{
  958. font-size: 0.86rem !important;
  959. font-weight: 500 !important;
  960. color: rgba(255,255,255,0.82) !important;
  961. }}
  962. /* Anchor targets. scroll-margin-top keeps the section header clear of
  963. Streamlit's fixed toolbar when a nav link jumps to it. */
  964. .entry-section {{
  965. scroll-margin-top: 4.5rem;
  966. }}
  967. /* Trailing runway for the last section (Sequence). An anchor can only scroll
  968. to the top of the viewport if enough content follows it; without this the
  969. final nav link lands short and the section never reaches the top, which
  970. reads as a broken link. 75vh clears it for the shortest realistic Sequence
  971. block without leaving a full blank screen below every entry. */
  972. .entry-tail {{
  973. min-height: 75vh;
  974. pointer-events: none;
  975. }}
  976. /* Entry hero: gene / accession line + key-fact strip */
  977. .entry-hero {{
  978. background: rgba(255,255,255,0.05);
  979. backdrop-filter: blur(14px);
  980. border: 1px solid rgba(116,162,183,0.25);
  981. border-left: 3px solid {COLORS['unreviewed']};
  982. border-radius: 12px;
  983. padding: 0.9rem 1.2rem 1rem 1.2rem;
  984. margin-bottom: 0.75rem;
  985. }}
  986. .entry-hero.is-swiss {{
  987. border-left-color: {COLORS['swiss_prot']};
  988. }}
  989. .entry-hero-title {{
  990. display: flex;
  991. align-items: center;
  992. gap: 0.7rem;
  993. flex-wrap: wrap;
  994. margin-bottom: 0.15rem;
  995. }}
  996. .entry-hero-title .entry-gene {{
  997. font-size: 1.6rem;
  998. font-weight: 700;
  999. color: #ffffff !important;
  1000. line-height: 1.2;
  1001. }}
  1002. .entry-hero-sub {{
  1003. font-family: monospace;
  1004. font-size: 0.82rem;
  1005. color: rgba(255,255,255,0.48) !important;
  1006. margin-bottom: 0.8rem;
  1007. }}
  1008. .entry-facts {{
  1009. display: flex;
  1010. flex-wrap: wrap;
  1011. gap: 0.4rem 2.2rem;
  1012. }}
  1013. .entry-facts > div {{ min-width: 7rem; }}
  1014. /* ── Password screen ── */
  1015. .login-card {{
  1016. max-width: 420px;
  1017. margin: 8vh auto;
  1018. background: rgba(255,255,255,0.06);
  1019. backdrop-filter: blur(16px);
  1020. border: 1px solid rgba(116,162,183,0.3);
  1021. border-radius: 18px;
  1022. padding: 2.5rem 2rem;
  1023. box-shadow: 0 8px 40px rgba(0,0,0,0.35);
  1024. text-align: center;
  1025. }}
  1026. .login-card h2 {{
  1027. background: linear-gradient(135deg, {COLORS['swiss_prot']}, {COLORS['unreviewed']});
  1028. -webkit-background-clip: text;
  1029. -webkit-text-fill-color: transparent;
  1030. background-clip: text;
  1031. font-size: 1.5rem;
  1032. margin-bottom: 0.3rem;
  1033. }}
  1034. .login-card p {{
  1035. color: rgba(255,255,255,0.6) !important;
  1036. font-size: 0.85rem;
  1037. margin-bottom: 1.2rem;
  1038. }}
  1039. /* ── Download button gradient ── */
  1040. .stDownloadButton > button {{
  1041. background: linear-gradient(135deg, {COLORS['swiss_prot']}, {COLORS['unreviewed']}) !important;
  1042. color: #fff !important;
  1043. border: none !important;
  1044. border-radius: 10px !important;
  1045. font-weight: 600 !important;
  1046. padding: 0.55rem 1.2rem !important;
  1047. transition: all 0.3s ease !important;
  1048. }}
  1049. .stDownloadButton > button:hover {{
  1050. background: linear-gradient(135deg, {COLORS['unreviewed']}, {COLORS['swiss_prot']}) !important;
  1051. transform: translateY(-2px) !important;
  1052. box-shadow: 0 6px 18px rgba(237,134,81,0.35) !important;
  1053. }}
  1054. /* ── Expander glass ── */
  1055. .streamlit-expanderHeader {{
  1056. background: rgba(255,255,255,0.08) !important;
  1057. backdrop-filter: blur(10px) !important;
  1058. border: 1px solid rgba(116,162,183,0.2) !important;
  1059. border-radius: 10px !important;
  1060. }}
  1061. .streamlit-expanderContent {{
  1062. background: rgba(255,255,255,0.04) !important;
  1063. backdrop-filter: blur(10px) !important;
  1064. border: 1px solid rgba(116,162,183,0.15) !important;
  1065. border-radius: 0 0 10px 10px !important;
  1066. }}
  1067. /* ── Metric containers (Streamlit native) ── */
  1068. [data-testid="metric-container"] {{
  1069. background: rgba(255,255,255,0.06) !important;
  1070. backdrop-filter: blur(8px) !important;
  1071. border: 1px solid rgba(116,162,183,0.18) !important;
  1072. border-radius: 10px !important;
  1073. }}
  1074. /* ── Dataframe glass wrapper ── */
  1075. .stDataFrame {{
  1076. background: rgba(255,255,255,0.04) !important;
  1077. backdrop-filter: blur(10px) !important;
  1078. border-radius: 12px 12px 12px 12px !important;
  1079. border: 1px solid rgba(116,162,183,0.18) !important;
  1080. }}
  1081. /* ── "How do I open an entry?" signpost above the table. The grid's row-select
  1082. box is the only in-place control it offers, and unlabelled it reads as
  1083. "select", so this points at it. ── */
  1084. .open-hint {{
  1085. display: flex;
  1086. align-items: center;
  1087. gap: 0.6rem;
  1088. background: rgba(237,134,81,0.10);
  1089. border: 1px solid rgba(237,134,81,0.40);
  1090. border-radius: 10px;
  1091. padding: 0.5rem 0.9rem;
  1092. margin: 0.2rem 0 0.5rem 0;
  1093. font-size: 0.88rem;
  1094. color: rgba(255,255,255,0.88) !important;
  1095. }}
  1096. .open-hint-arrow {{
  1097. color: {COLORS['unreviewed']};
  1098. font-size: 1rem;
  1099. line-height: 1;
  1100. }}
  1101. .open-hint strong {{ color: #ffffff !important; }}
  1102. /* ── Row-click prompt below the table (the detail card it used to dock here
  1103. now lives on its own entry page — see _render_entry_page) ── */
  1104. .detail-panel {{
  1105. background: rgba(255,255,255,0.04);
  1106. backdrop-filter: blur(14px);
  1107. border: 1px solid rgba(116,162,183,0.18);
  1108. border-top: 1px solid rgba(116,162,183,0.35);
  1109. border-radius: 0 0 12px 12px;
  1110. padding: 1rem 1.2rem 0.8rem 1.2rem;
  1111. margin-top: -1rem;
  1112. margin-bottom: 1rem;
  1113. position: relative;
  1114. }}
  1115. .detail-panel::before {{
  1116. content: '';
  1117. position: absolute;
  1118. top: 0; left: 5%; right: 5%; height: 1px;
  1119. background: linear-gradient(90deg, transparent, rgba(116,162,183,0.5), transparent);
  1120. }}
  1121. .detail-header {{
  1122. display: flex;
  1123. align-items: center;
  1124. gap: 0.7rem;
  1125. margin-bottom: 0.6rem;
  1126. padding-bottom: 0.5rem;
  1127. }}
  1128. .detail-header .detail-label {{
  1129. font-size: 1.15rem;
  1130. font-weight: 700;
  1131. color: #ffffff;
  1132. }}
  1133. .detail-header .detail-badge {{
  1134. margin-left: auto;
  1135. }}
  1136. .detail-prompt {{
  1137. text-align: center;
  1138. padding: 0.6rem 0;
  1139. font-size: 0.85rem;
  1140. color: rgba(255,255,255,0.4);
  1141. font-style: italic;
  1142. }}
  1143. /* ── Inputs dark navy ── */
  1144. .stTextInput > div > div > input,
  1145. .stNumberInput input,
  1146. input[type="number"],
  1147. input[type="text"] {{
  1148. background: #1a212d !important;
  1149. color: #ffffff !important;
  1150. border: 1px solid rgba(116,162,183,0.45) !important;
  1151. border-radius: 8px !important;
  1152. }}
  1153. .stMultiSelect > div > div {{
  1154. background: #1a212d !important;
  1155. border: 1px solid rgba(116,162,183,0.45) !important;
  1156. border-radius: 8px !important;
  1157. }}
  1158. .stButton > button {{
  1159. background: #1a212d !important;
  1160. color: #ffffff !important;
  1161. border: 1px solid rgba(116,162,183,0.45) !important;
  1162. border-radius: 8px !important;
  1163. transition: all 0.3s ease !important;
  1164. }}
  1165. .stButton > button:hover {{
  1166. background: #394254 !important;
  1167. border: 1px solid {COLORS['swiss_prot']} !important;
  1168. transform: translateY(-1px) !important;
  1169. box-shadow: 0 4px 12px rgba(116,162,183,0.25) !important;
  1170. }}
  1171. /* ── Transparent blocks ── */
  1172. .element-container, div[data-testid="stVerticalBlock"],
  1173. div[data-testid="stHorizontalBlock"], section[data-testid="stSidebar"] > div {{
  1174. background: transparent !important;
  1175. }}
  1176. /* ── Legend container ── */
  1177. .legend-container {{
  1178. background: rgba(26,32,44,0.75) !important;
  1179. border-radius: 12px !important;
  1180. padding: 0.9rem !important;
  1181. margin-top: 0.8rem !important;
  1182. border: 1px solid rgba(116,162,183,0.3) !important;
  1183. backdrop-filter: blur(8px) !important;
  1184. }}
  1185. /* ── Detail panel tabs ── */
  1186. [data-baseweb="tab-list"] {{
  1187. background: rgba(255,255,255,0.06) !important;
  1188. border-radius: 10px !important;
  1189. padding: 4px !important;
  1190. gap: 4px !important;
  1191. border: 1px solid rgba(116,162,183,0.25) !important;
  1192. }}
  1193. [data-baseweb="tab"] {{
  1194. background: transparent !important;
  1195. color: rgba(255,255,255,0.65) !important;
  1196. border-radius: 8px !important;
  1197. font-size: 0.92rem !important;
  1198. font-weight: 600 !important;
  1199. letter-spacing: 0.02em !important;
  1200. padding: 0.5rem 1.6rem !important;
  1201. border: 1px solid transparent !important;
  1202. transition: all 0.2s ease !important;
  1203. }}
  1204. [data-baseweb="tab"]:hover {{
  1205. background: rgba(116,162,183,0.15) !important;
  1206. color: rgba(255,255,255,0.9) !important;
  1207. border-color: rgba(116,162,183,0.3) !important;
  1208. }}
  1209. [aria-selected="true"][data-baseweb="tab"] {{
  1210. background: linear-gradient(135deg, rgba(116,162,183,0.30), rgba(237,134,81,0.20)) !important;
  1211. color: #ffffff !important;
  1212. border-color: rgba(116,162,183,0.5) !important;
  1213. box-shadow: 0 2px 8px rgba(0,0,0,0.25) !important;
  1214. }}
  1215. [data-baseweb="tab-highlight"] {{
  1216. display: none !important;
  1217. }}
  1218. [data-baseweb="tab-border"] {{
  1219. display: none !important;
  1220. }}
  1221. /* ── Sidebar section headers ── */
  1222. .sidebar-section-header {{
  1223. font-size: 0.68rem;
  1224. font-weight: 700;
  1225. text-transform: uppercase;
  1226. letter-spacing: 0.10em;
  1227. color: rgba(116,162,183,0.9) !important;
  1228. border-left: 2px solid rgba(116,162,183,0.55);
  1229. padding: 0.1rem 0 0.1rem 0.55rem;
  1230. margin: 0.9rem 0 0.45rem 0;
  1231. display: block;
  1232. }}
  1233. /* ── Table row hover & selection ── */
  1234. [data-testid="stDataFrame"] tr:hover td,
  1235. [data-testid="stDataFrame"] tr:hover th {{
  1236. background: rgba(116,162,183,0.08) !important;
  1237. transition: background 0.15s;
  1238. cursor: pointer;
  1239. }}
  1240. [data-testid="stDataFrame"] tr[aria-selected="true"] td {{
  1241. background: rgba(116,162,183,0.15) !important;
  1242. border-left: 2px solid rgba(116,162,183,0.65) !important;
  1243. }}
  1244. /* ── Active-filter summary bar ── */
  1245. .filter-bar {{
  1246. display: flex;
  1247. flex-wrap: wrap;
  1248. align-items: center;
  1249. gap: 0.4rem;
  1250. background: rgba(26,32,44,0.55);
  1251. border: 1px solid rgba(116,162,183,0.22);
  1252. border-radius: 12px;
  1253. padding: 0.55rem 0.8rem;
  1254. margin: -0.3rem 0 0.55rem 0;
  1255. }}
  1256. .filter-bar-title {{
  1257. font-size: 0.7rem;
  1258. text-transform: uppercase;
  1259. letter-spacing: 0.07em;
  1260. font-weight: 700;
  1261. color: rgba(116,162,183,0.95);
  1262. margin-right: 0.25rem;
  1263. }}
  1264. .filter-chip {{
  1265. display: inline-flex;
  1266. align-items: stretch;
  1267. overflow: hidden;
  1268. border-radius: 999px;
  1269. border: 1px solid rgba(116,162,183,0.4);
  1270. font-size: 0.76rem;
  1271. line-height: 1.6;
  1272. }}
  1273. .filter-chip-key {{
  1274. background: rgba(116,162,183,0.28);
  1275. color: #d5e6ef !important;
  1276. padding: 0 8px;
  1277. font-weight: 600;
  1278. }}
  1279. .filter-chip-val {{
  1280. background: rgba(237,134,81,0.16);
  1281. color: #f5d3bd !important;
  1282. padding: 0 9px;
  1283. font-weight: 500;
  1284. }}
  1285. .filter-bar-none {{
  1286. font-size: 0.8rem;
  1287. color: rgba(255,255,255,0.55);
  1288. font-style: italic;
  1289. }}
  1290. </style>
  1291. """, unsafe_allow_html=True)
  1292. # =============================================================================
  1293. # DATA LOADING — single unified source (the MASTER dataset)
  1294. # =============================================================================
  1295. # The dashboard derives every analysis view from one file,
  1296. # Code/data/microprotein_master.csv, by applying the canonical gold-standard
  1297. # filter and reproducing the column selections used by the
  1298. # Results/*/*_summary.csv generator scripts. The only inputs NOT contained in
  1299. # the master are the scRNA enrichment table (per gene x cluster) and the
  1300. # per-peptide tryptic / PROSIT-tier files, which are still read from disk.
  1301. MASTER_CSV = Path(__file__).resolve().parent.parent / "Code" / "data" / "microprotein_master.csv"
  1302. # Per-sequence P-site frame stats (Code/RP3_analysis/Psite_frame_mapping.py).
  1303. PSITE_FRAME_CSV = Path(__file__).resolve().parent / "RP3" / "Psite_frame_by_sequence.csv"
  1304. PSITE_FRAME_COLS = ["Psites_pct_frame0", "Psites_pct_frame1", "Psites_pct_frame2",
  1305. "Psites_frame0_RPKM"]
  1306. def _resolve_master_source(csv_path=MASTER_CSV):
  1307. """Return a path pandas can read for the master table.
  1308. The 88 MB `microprotein_master.csv` is git-ignored, so it is absent when
  1309. the repo is checked out on Streamlit Community Cloud. The compressed
  1310. `microprotein_master.zip` (a single-member archive) IS committed, and
  1311. pandas reads it directly via compression inference. Prefer the plain CSV
  1312. when present (local dev); otherwise fall back to the shipped zip.
  1313. """
  1314. p = Path(csv_path)
  1315. if p.exists():
  1316. return p
  1317. # Second name is the legacy (misspelled) filename, kept for backward compat.
  1318. for _zip_name in ("microprotein_master.zip", "micrprotein_master.zip"):
  1319. z = p.with_name(_zip_name)
  1320. if z.exists():
  1321. return z
  1322. return p # let the caller surface a clear FileNotFoundError
  1323. # UniProt curation annotation score (1–5) keyed by accession. Sourced from the
  1324. # UniProt proteome export (columns Entry, Annotation). Attached only to
  1325. # microproteins that ARE genuine UniProt entries (reviewed Swiss-Prot-MP or TrEMBL).
  1326. UNIPROT_ANNOT_TSV = Path(__file__).resolve().parent.parent / "Code" / "data" / "uniprotkb_proteome_UP000005640_2026_07_13.tsv"
  1327. @st.cache_data(show_spinner=False)
  1328. def load_uniprot_annotation_scores(_tsv_path=str(UNIPROT_ANNOT_TSV)):
  1329. """Map UniProt accession -> annotation score (1–5) from the proteome export.
  1330. Returns an empty dict if the TSV is unavailable so the dashboard degrades
  1331. gracefully (the annotation-score column simply stays blank).
  1332. """
  1333. p = Path(_tsv_path)
  1334. if not p.exists():
  1335. return {}
  1336. try:
  1337. ann = pd.read_csv(p, sep='\t', usecols=['Entry', 'Annotation'])
  1338. except Exception:
  1339. return {}
  1340. scores = pd.to_numeric(ann['Annotation'], errors='coerce')
  1341. return {str(e): float(s) for e, s in zip(ann['Entry'], scores) if pd.notna(s)}
  1342. # Per-microprotein rescue for the N-terminus filter: does at least one of its
  1343. # N-terminal (aa<=2) tryptic peptides carry 1-2 amino-acid substitutions vs.
  1344. # its matched UniProt isoform? See
  1345. # Code/Microprotein_annotation_summary/compute_nterm_peptide_substitutions.py.
  1346. NTERM_SUBSTITUTIONS_CSV = Path(__file__).resolve().parent.parent / "Code" / "data" / "blast_nterm_peptide_substitutions.csv"
  1347. @st.cache_data(show_spinner=False)
  1348. def load_nterm_substitutions(_csv_path=str(NTERM_SUBSTITUTIONS_CSV)):
  1349. """Map microprotein sequence -> (has_rescue, detail rows).
  1350. `has_rescue` is True when any N-terminal peptide has a comparable
  1351. alignment (coverage>=80%) to its matched isoform with 1-2 substitutions --
  1352. direct evidence that peptide isn't identical to the canonical protein, so
  1353. it can't be a bare Met-excision fragment of it. `detail` carries the
  1354. per-peptide rows for display in the microprotein detail view.
  1355. Returns an empty dict if the file is unavailable so the N-terminus filter
  1356. degrades gracefully to its unadjusted (stricter) behavior.
  1357. """
  1358. p = Path(_csv_path)
  1359. if not p.exists():
  1360. return {}
  1361. try:
  1362. sub = pd.read_csv(p)
  1363. except Exception:
  1364. return {}
  1365. result = {}
  1366. for seq, group in sub.groupby('sequence'):
  1367. rescues = group[group['comparable'].astype(bool)
  1368. & group['mismatch_count'].between(1, 2)]
  1369. result[seq] = {
  1370. 'has_rescue': not rescues.empty,
  1371. 'detail': group.to_dict('records'),
  1372. }
  1373. return result
  1374. @st.cache_data(show_spinner=False)
  1375. def load_decay_classes(_tsv_path=str(DECAY_CLASS_TSV)):
  1376. """Map master gene_id -> ORF-rules label ('NMD · mORF Intact', ...).
  1377. Keyed on the assignments TSV's `smorf_id`, which is the same identifier as
  1378. the master's `gene_id` (the master has no `orf_id` column). The join is
  1379. exact and 1:1; the ~867 `_alt_initiation_N` proteoforms in the TSV have no
  1380. master counterpart and are deliberately left unmatched -- collapsing them
  1381. onto their parent ORF would assign conflicting classes for 9 parents.
  1382. The TSV's `category` spans six values: the NMD x main-ORF-disruption 2x2
  1383. (dark_green/light_green/gold/firebrick) plus the two non-stop-decay
  1384. classes (nsd_disrupted/nsd_intact), which the generator scores per smORF
  1385. ahead of that 2x2 and which therefore override it. Categories absent from
  1386. DECAY_CLASS_FROM_CATEGORY are dropped and fall through to 'NA', so a new
  1387. upstream category must be added there or it silently reads as unassessed.
  1388. Returns an empty dict if the file is unavailable so the column degrades
  1389. gracefully to all-'NA' (e.g. on Streamlit Cloud / the HF Space).
  1390. """
  1391. p = Path(_tsv_path)
  1392. if not p.exists():
  1393. return {}
  1394. try:
  1395. a = pd.read_csv(p, sep='\t', usecols=['smorf_id', 'category'])
  1396. except Exception:
  1397. return {}
  1398. return {r.smorf_id: DECAY_CLASS_FROM_CATEGORY[r.category]
  1399. for r in a.itertuples()
  1400. if r.category in DECAY_CLASS_FROM_CATEGORY}
  1401. # Best-tier ranking used to collapse the bracketed PROSIT `Confidence` list.
  1402. _CONFIDENCE_RANK = {'Strong': 0, 'Moderate': 1, 'Weak': 2, 'Insufficient': 3}
  1403. def _best_confidence(val):
  1404. """Pick the best (lowest-rank) PROSIT tier. Accepts either a plain scalar tier
  1405. string (current master) or a bracketed list literal of tiers (legacy master)."""
  1406. if pd.isna(val):
  1407. return None
  1408. text = str(val).strip()
  1409. if text in _CONFIDENCE_RANK: # plain scalar tier (current master format)
  1410. return text
  1411. try:
  1412. items = ast.literal_eval(text)
  1413. except (ValueError, SyntaxError):
  1414. return None
  1415. if not isinstance(items, list):
  1416. return None
  1417. valid = [x for x in items if x in _CONFIDENCE_RANK]
  1418. if not valid:
  1419. return None
  1420. return min(valid, key=lambda x: _CONFIDENCE_RANK[x])
  1421. @st.cache_data(show_spinner=False)
  1422. def load_and_filter_master(_csv_path=str(MASTER_CSV)):
  1423. """Apply the canonical gold-standard filter to the MASTER dataset.
  1424. This is an inlined copy of
  1425. Code/gold_standard_filtering_criteria.load_and_filter_master — the
  1426. authoritative filtering pipeline shared by all Results-summary generator
  1427. scripts. Keep the two in sync if the criteria ever change.
  1428. """
  1429. df = pd.read_csv(_resolve_master_source(_csv_path), low_memory=False)
  1430. # Step 1: drop contaminant / non-human entries
  1431. if "gene_symbol" in df.columns:
  1432. df = df[~df["gene_symbol"].str.contains(r"SHEEP|Peptide|Horse|BOVIN",
  1433. case=False, na=False)]
  1434. # Step 2: keep relevant databases
  1435. df = df[df["Database"].str.contains("TrEMBL|Salk|Swiss-Prot", na=False)].copy()
  1436. # Step 3: remap Database labels (TrEMBL is Salk-derived; Swiss-Prot -> Swiss-Prot-MP)
  1437. df.loc[df["Database"] == "TrEMBL", "smorf_type"] = "TrEMBL"
  1438. df["Database"] = df["Database"].replace({
  1439. "TrEMBL": "Salk",
  1440. "Swiss-Prot": "Swiss-Prot-MP",
  1441. })
  1442. # Step 4: evidence flags
  1443. ribocode_pass = df["RiboCode"].astype(str).str.strip().str.upper() == "TRUE" # noqa: E712
  1444. df["has_RiboSAM"] = (
  1445. ribocode_pass
  1446. & df["shortstop_label"].isin(["SAM-Secreted", "SAM-Intracellular"])
  1447. & (~df["smorf_type"].isin(["iORF"]))
  1448. & (~df["smorf_type"].str.contains("iso", case=False, na=False))
  1449. )
  1450. df["has_LooseRiboSAM"] = (
  1451. ribocode_pass
  1452. & df["shortstop_label"].isin(["SAM-Secreted", "SAM-Intracellular"])
  1453. & (~df["smorf_type"].isin(["Iso"]))
  1454. )
  1455. df["DDA_evidence"] = df["total_unique_spectral_counts"] > 0
  1456. df["DIA_evidence"] = (
  1457. df["has_LooseRiboSAM"]
  1458. & df["Global.PG.Q.Value"].notna()
  1459. & (df["Global.PG.Q.Value"] <= 0.01)
  1460. & (df["Proteotypic"] == 1)
  1461. )
  1462. df["DIA_expression"] = (
  1463. df["Global.PG.Q.Value"].notna()
  1464. & (df["Global.PG.Q.Value"] <= 0.01)
  1465. & (df["Proteotypic"] == 1)
  1466. )
  1467. # DIA proteomics is intentionally excluded from this dashboard: only DDA
  1468. # mass-spec counts as MS evidence (DIA_evidence/DIA_expression columns are
  1469. # retained for reference but are NOT used for inclusion or labeling).
  1470. df["has_MS"] = df["DDA_evidence"]
  1471. # Coerce numeric-evidence columns that may contain stray non-numeric
  1472. # tokens (e.g. "FALSE") in the upstream master
  1473. for _col in ("RP3_Default", "total_razor_spectral_counts"):
  1474. if _col in df.columns:
  1475. df[_col] = pd.to_numeric(df[_col], errors="coerce").fillna(0)
  1476. # Step 5: microproteins only
  1477. df = df[df["protein_class_length"] == "Microprotein"].copy()
  1478. # Step 6: split by database, apply evidence filters, dedup by sequence
  1479. mp_swiss = (
  1480. df[df["Database"] == "Swiss-Prot-MP"]
  1481. .loc[lambda d:
  1482. (d["RP3_Default"] > 0)
  1483. | (d["total_razor_spectral_counts"] > 0)]
  1484. .drop_duplicates(subset="sequence")
  1485. )
  1486. mp_salk = (
  1487. df[(df["Database"] == "Salk")
  1488. & (df["has_RiboSAM"] | df["DDA_evidence"])]
  1489. .sort_values("has_MS", ascending=False)
  1490. .drop_duplicates(subset="sequence")
  1491. )
  1492. mp = pd.concat([mp_swiss, mp_salk], ignore_index=True)
  1493. # Step 7: attach the smORF decay class. Must happen HERE, while gene_id is
  1494. # still around -- the per-analysis merge downstream selects an explicit
  1495. # column allowlist keyed on `sequence` and drops gene_id entirely.
  1496. _decay_map = load_decay_classes()
  1497. if _decay_map and "gene_id" in mp.columns:
  1498. mp["NMD_Decay_Class"] = mp["gene_id"].map(_decay_map).fillna("NA")
  1499. else:
  1500. mp["NMD_Decay_Class"] = "NA"
  1501. return mp
  1502. # --- Per-analysis derivations (mirror the Results-summary generator scripts) ---
  1503. def _derive_annotation(mp):
  1504. """Salk discovery summary (Annotation Status, MS Evidence Type, DDA Grade)."""
  1505. df = mp[mp["Database"] == "Salk"].copy()
  1506. # Evidence Type matches Figure 7 Panel J: RiboCode-ShortStop = RiboSAM
  1507. # translation support without DDA mass-spec detection (has_RiboSAM & ~DDA_evidence);
  1508. # everything else (DDA-detected) is MS. DIA is not counted in this dashboard.
  1509. df["Annotation Status"] = (df["has_RiboSAM"] & ~df["DDA_evidence"]).map(
  1510. {True: "RiboCode-ShortStop", False: "MS"})
  1511. df["MS_type"] = df["DDA_evidence"].map({True: "DDA", False: "No MS"})
  1512. df["DDA Grade"] = df["Confidence"].apply(_best_confidence)
  1513. df.loc[df["DDA Grade"].isna() & (df["Annotation Status"] == "RiboCode-ShortStop"), "DDA Grade"] = "No MS"
  1514. df.loc[df["DDA Grade"].isna() & (df["Annotation Status"] == "MS"), "DDA Grade"] = "No PROSIT"
  1515. # Kozak context strength + non-AUG start codon assessment — Salk-only columns
  1516. # (Code/data/microprotein_master.csv nonATG_*/kozak_* fields), passed through
  1517. # unrenamed so extract_unified_fields can pick them up by suffix.
  1518. _nonatg_kozak_cols = [c for c in (
  1519. "kozak_full_kozak_strength", "kozak_downstream_kozak_strength",
  1520. "kozak_kozak_class_canonical", "kozak_kozak_window",
  1521. "nonATG_annotated_start_codon", "nonATG_annotated_start_is_initiator",
  1522. "nonATG_has_supported_initiation_site",
  1523. "nonATG_annotated_context_strength", "nonATG_n_initiation_sites",
  1524. "nonATG_no_site_reason", "nonATG_site_type",
  1525. "nonATG_initiation_codon_position", "nonATG_initiation_codon",
  1526. "nonATG_cognate_status", "nonATG_initiation_tier_name",
  1527. "nonATG_initiator_aa", "nonATG_predicted_sequence",
  1528. ) if c in df.columns]
  1529. summary = df[[
  1530. "sequence", "CLICK_UCSC", "genomic_coordinates", "smorf_type", "gene_name",
  1531. "protein_length", "Annotation Status", "MS_type", "DDA Grade", "mean_phylocsf",
  1532. *_nonatg_kozak_cols,
  1533. ]].rename(columns={
  1534. "genomic_coordinates": "smORF Coordinates",
  1535. "smorf_type": "smORF Class",
  1536. "gene_name": "Parent Gene",
  1537. "protein_length": "Microprotein Length",
  1538. "sequence": "Microprotein Sequence",
  1539. "MS_type": "MS Evidence Type",
  1540. "mean_phylocsf": "PhyloCSF Score",
  1541. })
  1542. summary["_starts_M"] = summary["Microprotein Sequence"].str.startswith("M")
  1543. summary = summary.sort_values(by="_starts_M", ascending=False).drop(columns="_starts_M")
  1544. return summary
  1545. def _derive_proteomics(mp):
  1546. return mp[[
  1547. "sequence", "CLICK_UCSC",
  1548. "TMT_log2fc_50pct_missing", "TMT_t_statistic_50pct_missing", "TMT_df_50pct_missing",
  1549. "TMT_conf_low_50pct_missing", "TMT_conf_high_50pct_missing", "TMT_cohens_d_50pct_missing",
  1550. "TMT_pvalue_50pct_missing", "TMT_qvalue_50pct_missing",
  1551. "TMT_log2fc_0pct_missing", "TMT_t_statistic_0pct_missing", "TMT_df_0pct_missing",
  1552. "TMT_conf_low_0pct_missing", "TMT_conf_high_0pct_missing", "TMT_cohens_d_0pct_missing",
  1553. "TMT_pvalue_0pct_missing", "TMT_qvalue_0pct_missing",
  1554. "rate_control", "rate_ad", "Database",
  1555. "protein_class_length", "gene_symbol", "gene_name", "protein_length",
  1556. "start_codon", "smorf_type", "total_razor_spectral_counts",
  1557. "total_unique_spectral_counts", "mean_phylocsf",
  1558. "blastp_uniprot_accession_match", "microprotein_percentage_match",
  1559. "blastp_alignment_length", "evalue", "blastp_bit",
  1560. ]].copy()
  1561. def _derive_rp3(mp):
  1562. df = mp[mp["Database"].isin(["Salk", "Swiss-Prot-MP"])][[
  1563. "sequence", "CLICK_UCSC", "RP3_Default", "RP3_MM_Amb", "RP3_Amb", "RP3_MM",
  1564. "RiboCode", "Database", "gene_name", "gene_symbol", "protein_length",
  1565. "start_codon", "smorf_type", "total_unique_spectral_counts",
  1566. "total_razor_spectral_counts", "mean_phylocsf",
  1567. ]].copy()
  1568. if PSITE_FRAME_CSV.exists():
  1569. psite = pd.read_csv(PSITE_FRAME_CSV, usecols=["sequence"] + PSITE_FRAME_COLS)
  1570. df = df.merge(psite, on="sequence", how="left")
  1571. return df
  1572. def _derive_shortread(mp):
  1573. return mp[[
  1574. "sequence", "CLICK_UCSC", "gene_symbol", "genomic_coordinates",
  1575. "rosmapRNA_baseMean", "rosmapRNA_log2FoldChange", "rosmapRNA_pvalue",
  1576. "rosmapRNA_padj", "rosmapRNA_non_smorf_hit", "rosmapRNA_body",
  1577. "ROSMAP_BulkRNAseq_CPM", "correlation_mainORF_nonAD_rosmap",
  1578. "correlation_mainORF_AD_rosmap", "rosmap_lrt_additive_p",
  1579. "rosmap_lrt_interaction_p", "msbbRNA_baseMean", "msbbRNA_log2FoldChange",
  1580. "msbbRNA_pvalue", "msbbRNA_padj", "msbbRNA_non_smorf_hit", "msbbRNA_body",
  1581. "MSBB_BulkRNAseq_CPM", "correlation_mainORF_nonAD_msbb",
  1582. "correlation_mainORF_AD_msbb", "Database", "gene_name", "protein_length",
  1583. "start_codon", "smorf_type", "total_razor_spectral_counts",
  1584. "total_unique_spectral_counts", "mean_phylocsf",
  1585. ]].copy()
  1586. def _derive_longread(mp):
  1587. df = mp[[
  1588. "sequence", "CLICK_UCSC", "nanopore_baseMean", "nanopore_log2FoldChange",
  1589. "nanopore_pvalue", "nanopore_padj", "Database", "gene_name", "gene_symbol",
  1590. "protein_length", "start_codon", "smorf_type", "total_razor_spectral_counts",
  1591. "total_unique_spectral_counts",
  1592. ]].copy()
  1593. return df[(df["Database"] == "Salk") | df["nanopore_baseMean"].notna()]
  1594. def _derive_shortstop(mp):
  1595. df = mp[mp["Database"] == "Salk"].copy()
  1596. df["Annotation Status"] = (df["has_RiboSAM"] & ~df["DDA_evidence"]).map(
  1597. {True: "RiboCode-ShortStop", False: "MS"})
  1598. return df.loc[df["shortstop_label"].notna(), [
  1599. "sequence", "CLICK_UCSC", "gene_symbol", "gene_name", "smorf_type",
  1600. "shortstop_label", "shortstop_score", "Annotation Status", "mean_phylocsf",
  1601. ]].rename(columns={
  1602. "sequence": "Microprotein Sequence",
  1603. "gene_symbol": "smORF ID",
  1604. "gene_name": "Gene Body (Name)",
  1605. "smorf_type": "smORF Class",
  1606. "shortstop_label": "ShortStop Label",
  1607. "shortstop_score": "ShortStop Score",
  1608. "mean_phylocsf": "PhyloCSF Score",
  1609. })
  1610. def load_analysis_results():
  1611. """Build every analysis view from the single master file.
  1612. Returns an ordered dict {analysis_name: {"df": DataFrame, "type": str}}.
  1613. All views except scRNA Enrichment are derived in-memory from the
  1614. gold-standard-filtered master; scRNA is read from its own table because it
  1615. is not contained in the master dataset.
  1616. """
  1617. mp = load_and_filter_master()
  1618. scrna_path = SCRNA_SUMMARY_FILE
  1619. scrna_df = None
  1620. if INCLUDE_SCRNA and scrna_path.exists():
  1621. try:
  1622. scrna_df = pd.read_csv(scrna_path, low_memory=False)
  1623. except Exception as e:
  1624. st.warning(f"Could not load scRNA Enrichment: {e}")
  1625. analyses = {
  1626. "Annotation Summary": {"df": _derive_annotation(mp), "type": "unreviewed_only"},
  1627. "Proteomics (TMT)": {"df": _derive_proteomics(mp), "type": "mixed"},
  1628. "Proteomics + RiboSeq (RP3)": {"df": _derive_rp3(mp), "type": "mixed"},
  1629. "Short-Read RNA in AD": {"df": _derive_shortread(mp), "type": "mixed"},
  1630. "Long-Read RNA in AD": {"df": _derive_longread(mp), "type": "unreviewed_only"},
  1631. "scRNA Enrichment": {"df": scrna_df, "type": "scrna"} if INCLUDE_SCRNA else None,
  1632. "ShortStop Classification": {"df": _derive_shortstop(mp), "type": "unreviewed_only"},
  1633. }
  1634. # Drop conditional/None entries and any view that failed to build.
  1635. return {k: v for k, v in analyses.items()
  1636. if v is not None and v.get("df") is not None}
  1637. @st.cache_data(show_spinner=False)
  1638. def load_and_merge_all_data():
  1639. """Merge every analysis view by sequence, anchored on the master dataset.
  1640. The tryptic-peptide, PROSIT and N-terminal-acetylation columns come straight
  1641. from the master (peptide_sequence/start/end -> Tryptic_*, Confidence for the
  1642. per-peptide PROSIT tiers, Nt_acetyl_* for the Nt-acetylated subset), so no
  1643. separate peptide CSV or tier-ID files are read.
  1644. """
  1645. analysis_files = load_analysis_results()
  1646. # Base frame: tryptic peptides + PROSIT confidence + Nt-acetylation straight
  1647. # from the master. The Nt_acetyl_* columns are ';'-delimited and parallel to
  1648. # each other (one entry per Nt-acetylated tryptic peptide), but are a *subset*
  1649. # of Tryptic_Peptides — they are not positionally aligned with it.
  1650. mp = load_and_filter_master()
  1651. _base_cols = [c for c in ['sequence', 'peptide_sequence', 'start', 'end',
  1652. 'Confidence', 'NMD_Decay_Class'] if c in mp.columns]
  1653. _acetyl_cols = [c for c in ('Nt_acetyl_tryptic_peptide', 'Nt_acetyl_N_PSMs',
  1654. 'Nt_acetyl_PSM_fraction') if c in mp.columns]
  1655. master_df = mp[_base_cols + _acetyl_cols].rename(
  1656. columns={
  1657. 'peptide_sequence': 'Tryptic_Peptides',
  1658. 'start': 'Tryptic_Start_Positions',
  1659. 'end': 'Tryptic_End_Positions',
  1660. 'Nt_acetyl_tryptic_peptide': 'Nt_Acetyl_Peptides',
  1661. 'Nt_acetyl_N_PSMs': 'Nt_Acetyl_N_PSMs',
  1662. 'Nt_acetyl_PSM_fraction': 'Nt_Acetyl_PSM_Fraction',
  1663. }
  1664. ).copy()
  1665. master_df['Tryptic_Peptides_present'] = master_df['Tryptic_Peptides'].notna()
  1666. if 'Nt_Acetyl_Peptides' in master_df.columns:
  1667. master_df['Nt_Acetylated'] = (
  1668. master_df['Nt_Acetyl_Peptides'].notna()
  1669. & master_df['Nt_Acetyl_Peptides'].astype(str).str.strip().ne('')
  1670. )
  1671. else:
  1672. master_df['Nt_Acetylated'] = False
  1673. for analysis_name, info in analysis_files.items():
  1674. df = info.get('df')
  1675. if df is not None and not df.empty:
  1676. try:
  1677. df = df.copy()
  1678. if 'Microprotein Sequence' in df.columns:
  1679. df = df.rename(columns={'Microprotein Sequence': 'sequence'})
  1680. elif 'sequence' not in df.columns:
  1681. continue
  1682. df[f'{analysis_name}_present'] = True
  1683. prefix = analysis_name.replace(' ', '_').replace('(', '').replace(')', '').replace('+', '_')
  1684. cols_to_rename = {}
  1685. for col in df.columns:
  1686. if col not in ['sequence', f'{analysis_name}_present']:
  1687. cols_to_rename[col] = f"{prefix}_{col}"
  1688. df = df.rename(columns=cols_to_rename)
  1689. if master_df.empty:
  1690. master_df = df.copy()
  1691. else:
  1692. master_df = pd.merge(master_df, df, on='sequence', how='outer', suffixes=('', '_dup'))
  1693. dup_cols = [c for c in master_df.columns if c.endswith('_dup')]
  1694. master_df = master_df.drop(columns=dup_cols)
  1695. except Exception as e:
  1696. st.warning(f"Could not load {analysis_name}: {e}")
  1697. presence_cols = [c for c in master_df.columns if c.endswith('_present')]
  1698. for col in presence_cols:
  1699. master_df[col] = master_df[col].fillna(False)
  1700. if not master_df.empty and 'sequence' in master_df.columns:
  1701. data_cols = [c for c in master_df.columns if c != 'sequence' and not c.endswith('_present')]
  1702. master_df['_completeness'] = master_df[data_cols].notna().sum(axis=1)
  1703. master_df = master_df.sort_values(['sequence', '_completeness'], ascending=[True, False])
  1704. master_df = master_df.drop_duplicates(subset=['sequence'], keep='first')
  1705. master_df = master_df.drop('_completeness', axis=1)
  1706. ann_col = [c for c in master_df.columns
  1707. if 'annotation' in c.lower() and 'summary' in c.lower() and c.endswith('_present')]
  1708. if ann_col:
  1709. ann_mask = master_df[ann_col[0]] == True
  1710. # Keep original unreviewed-anchored set, OR include Swiss-Prot microproteins.
  1711. db_cols = [c for c in master_df.columns if 'database' in c.lower()]
  1712. if db_cols:
  1713. swiss_mask = master_df[db_cols].apply(
  1714. lambda row: any('swiss-prot' in str(v).lower() for v in row if pd.notna(v)),
  1715. axis=1
  1716. )
  1717. else:
  1718. swiss_mask = pd.Series(False, index=master_df.index)
  1719. # Microprotein criterion: sequence length <= 151 aa.
  1720. seq_len = master_df['sequence'].astype(str).str.len() if 'sequence' in master_df.columns else pd.Series(999, index=master_df.index)
  1721. microprotein_mask = seq_len <= 151
  1722. # Only include Swiss-Prot entries that appear in the Proteomics CSV.
  1723. # The 752 Swiss-Prot-MP entries added after the first submission have no
  1724. # MS evidence and should not be shown in the dashboard.
  1725. prot_col = next((c for c in master_df.columns if c == 'Proteomics (TMT)_present'), None)
  1726. if prot_col:
  1727. prot_present_mask = master_df[prot_col].fillna(False).astype(bool)
  1728. else:
  1729. prot_present_mask = pd.Series(True, index=master_df.index)
  1730. master_df = master_df[ann_mask | (swiss_mask & microprotein_mask & prot_present_mask)]
  1731. return master_df
  1732. def _has_non_nterm_peptide(pos_str):
  1733. """True if any tryptic peptide starts at aa >= 3.
  1734. A peptide starting at aa 1 or 2 is exactly what Met excision of the *parent*
  1735. ORF would produce, so it is not on its own evidence for the microprotein;
  1736. a peptide starting at aa 3 or later is evidence independent of that artefact.
  1737. `pos_str` is the master's Python-list-literal of start positions.
  1738. """
  1739. if not pos_str or pd.isna(pos_str):
  1740. return False # no tryptic data is not evidence of anything
  1741. try:
  1742. positions = ast.literal_eval(str(pos_str))
  1743. except (ValueError, SyntaxError):
  1744. return False
  1745. if isinstance(positions, (int, float)):
  1746. positions = [positions]
  1747. try:
  1748. return any(int(p) >= 3 for p in positions)
  1749. except (TypeError, ValueError):
  1750. return False
  1751. def _non_nterm_series(df, index=None):
  1752. """Boolean `Has_Non_Nterm_Peptide` for `df`, derived on the fly if absent.
  1753. Normally a plain column lookup — the column is precomputed once per data
  1754. load. The fallback matters when a long-running process holds a dataframe
  1755. cached before the column existed: st.cache_data keys on the decorated
  1756. function's own bytecode, so editing a helper it calls does not invalidate
  1757. it. Recomputing there costs a parse but keeps the filter honest; silently
  1758. returning "no filter" would make a stale cache look like a broken widget.
  1759. """
  1760. if 'Has_Non_Nterm_Peptide' in df.columns:
  1761. return df['Has_Non_Nterm_Peptide'].fillna(False).astype(bool)
  1762. if 'Tryptic_Start_Positions' in df.columns:
  1763. return df['Tryptic_Start_Positions'].map(_has_non_nterm_peptide).astype(bool)
  1764. return None
  1765. # Evidence & Quality "≥N peptides" tiers (distinct tryptic peptide sequences).
  1766. # Nested like the significance tiers: ticking several is the same as ticking
  1767. # the lowest.
  1768. MIN_PEPTIDE_TIERS = [2, 3, 4]
  1769. def _count_peptides(pep_str):
  1770. """Number of distinct tryptic peptide sequences in the master's list literal."""
  1771. if not pep_str or pd.isna(pep_str):
  1772. return 0
  1773. try:
  1774. peps = ast.literal_eval(str(pep_str))
  1775. except (ValueError, SyntaxError):
  1776. peps = [p for p in str(pep_str).split(';') if p.strip()]
  1777. if isinstance(peps, str):
  1778. peps = [peps]
  1779. try:
  1780. return len({str(p).strip() for p in peps if str(p).strip()})
  1781. except TypeError:
  1782. return 0
  1783. def _peptide_count_series(df):
  1784. """`Tryptic_Peptide_Count` for `df`, derived on the fly if absent (same
  1785. stale-cache fallback as _non_nterm_series)."""
  1786. if 'Tryptic_Peptide_Count' in df.columns:
  1787. return df['Tryptic_Peptide_Count'].fillna(0).astype(int)
  1788. if 'Tryptic_Peptides' in df.columns:
  1789. return df['Tryptic_Peptides'].map(_count_peptides).astype(int)
  1790. return None
  1791. def extract_unified_fields(master_df):
  1792. """Extract and unify key fields from the merged dataset."""
  1793. ud = master_df.copy()
  1794. # Parent Gene
  1795. cols = [c for c in master_df.columns
  1796. if ('parent' in c.lower() and 'gene' in c.lower())
  1797. or c.lower().endswith('gene_name')
  1798. or c.lower().endswith('gene_symbol')
  1799. or (c.lower().endswith('_gene') and 'id' not in c.lower())
  1800. or c.lower().endswith('gene body (name)')]
  1801. if cols:
  1802. ud['Parent_Gene'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  1803. # smORF Class
  1804. cols = [c for c in master_df.columns
  1805. if ('smorf' in c.lower() and ('class' in c.lower() or 'type' in c.lower()))
  1806. and 'id' not in c.lower().split('smorf')[1]
  1807. and 'coordinates' not in c.lower()]
  1808. if cols:
  1809. ud['smORF_Class'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  1810. else:
  1811. ud['smORF_Class'] = pd.NA
  1812. # Swiss-Prot-MP entries have no smorf_type — assign class so they appear in group filters.
  1813. # (After the merge, the bare 'Database' column doesn't exist; use the prefixed cols.)
  1814. _db_cols_raw = [c for c in master_df.columns if 'database' in c.lower()]
  1815. if _db_cols_raw:
  1816. _is_swiss_mp = master_df[_db_cols_raw].apply(
  1817. lambda row: any('swiss-prot' in str(v).lower() for v in row if pd.notna(v)), axis=1
  1818. )
  1819. ud.loc[ud['smORF_Class'].isna() & _is_swiss_mp, 'smORF_Class'] = 'Swiss-Prot-MP'
  1820. # Display label (no remapping currently; hook for future overrides)
  1821. ud['smORF_Display'] = ud['smORF_Class'].map(
  1822. lambda v: SMORF_DISPLAY_LABEL.get(str(v), v) if pd.notna(v) else v
  1823. )
  1824. # Parent group (Upstream, Downstream, TrEMBL/AltORF, etc.)
  1825. ud['smORF_Group'] = ud['smORF_Class'].map(
  1826. lambda v: SMORF_CHILD_TO_PARENT.get(str(v), str(v)) if pd.notna(v) else v
  1827. )
  1828. # smORF decay class (NMD × main-ORF disruption). Already joined on gene_id
  1829. # back in load_and_filter_master and carried through the merge; this only
  1830. # normalizes rows the outer merge introduced.
  1831. # NB: the column is named NMD_* on purpose -- the smORF_Class discovery
  1832. # above pattern-matches any column containing 'smorf' + 'class'/'type', so
  1833. # a name like 'smorf_decay_class' would be silently bfilled into
  1834. # smORF_Class and shadow smorf_type.
  1835. if 'NMD_Decay_Class' in ud.columns:
  1836. ud['NMD_Decay_Class'] = ud['NMD_Decay_Class'].fillna('NA')
  1837. else:
  1838. ud['NMD_Decay_Class'] = 'NA'
  1839. # ShortStop Label
  1840. cols = [c for c in master_df.columns if 'shortstop' in c.lower() and 'label' in c.lower()]
  1841. if cols:
  1842. ud['ShortStop_Label'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  1843. # Annotation Status
  1844. cols = [c for c in master_df.columns if 'annotation' in c.lower() and 'status' in c.lower()]
  1845. if cols:
  1846. ud['Annotation_Status'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  1847. # Unique Spectral Counts
  1848. cols = [c for c in master_df.columns if 'unique_spectral_counts' in c.lower()]
  1849. if cols:
  1850. ud['Unique_Spectral_Counts'] = pd.to_numeric(
  1851. master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  1852. # Razor Spectral Counts
  1853. cols = [c for c in master_df.columns if 'razor_spectral_counts' in c.lower()]
  1854. if cols:
  1855. ud['Razor_Spectral_Counts'] = pd.to_numeric(
  1856. master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  1857. # UCSC Link
  1858. cols = [c for c in master_df.columns if 'ucsc' in c.lower() or 'CLICK_UCSC' in c]
  1859. if cols:
  1860. ud['UCSC_Link'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  1861. # Protein Length
  1862. cols = [c for c in master_df.columns if 'length' in c.lower() and 'class' not in c.lower()]
  1863. if cols:
  1864. ud['Protein_Length'] = pd.to_numeric(
  1865. master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  1866. # Fill missing Protein_Length from sequence length
  1867. if 'sequence' in ud.columns:
  1868. missing = ud['Protein_Length'].isna()
  1869. ud.loc[missing, 'Protein_Length'] = ud.loc[missing, 'sequence'].astype(str).str.len()
  1870. # Start Codon (the plain "ATG"/"nonATG" call). Excludes 'nonatg' columns —
  1871. # nonATG_annotated_start_codon also matches 'start'+'codon' but holds the raw
  1872. # triplet (e.g. "GCC"), not the coarse ATG/nonATG label.
  1873. cols = [c for c in master_df.columns
  1874. if 'start' in c.lower() and 'codon' in c.lower() and 'nonatg' not in c.lower()]
  1875. if cols:
  1876. ud['Start_Codon'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  1877. # ShortStop Score
  1878. cols = [c for c in master_df.columns if 'shortstop' in c.lower() and 'score' in c.lower()]
  1879. if cols:
  1880. ud['ShortStop_Score'] = pd.to_numeric(
  1881. master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  1882. # PhyloCSF Score
  1883. cols = [c for c in master_df.columns if 'phylocsf' in c.lower() or 'mean_phylocsf' in c.lower()]
  1884. if cols:
  1885. ud['PhyloCSF_Score'] = pd.to_numeric(
  1886. master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  1887. min_val = ud['PhyloCSF_Score'].min()
  1888. if pd.isna(min_val):
  1889. min_val = -1000.0
  1890. # Fill NaN with min to push missing-data rows to the bottom of score-based sorts
  1891. ud['PhyloCSF_Score'] = ud['PhyloCSF_Score'].fillna(min_val)
  1892. # Tryptic peptides (+ the Nt-acetylated subset of them)
  1893. for col in ['Tryptic_Peptides', 'Tryptic_Protein_ID', 'Tryptic_Start_Positions', 'Tryptic_End_Positions',
  1894. 'Nt_Acetyl_Peptides', 'Nt_Acetyl_N_PSMs', 'Nt_Acetyl_PSM_Fraction']:
  1895. if col in master_df.columns:
  1896. ud[col] = master_df[col]
  1897. # N-terminal acetylation flag + summary numbers. The master stores ';'-delimited
  1898. # parallel lists (one entry per Nt-acetylated peptide); collapse them to a total
  1899. # PSM count and the best (highest) per-peptide PSM fraction so both are sortable.
  1900. ud['Nt_Acetylated'] = (
  1901. master_df['Nt_Acetylated'].fillna(False).astype(bool)
  1902. if 'Nt_Acetylated' in master_df.columns
  1903. else pd.Series(False, index=ud.index)
  1904. )
  1905. # N-terminal-peptide substitution rescue (see load_nterm_substitutions) --
  1906. # False for the overwhelming majority of rows, which have no comparable
  1907. # 1-2-substitution N-terminal peptide at all.
  1908. _nterm_sub_map = load_nterm_substitutions()
  1909. ud['Nterm_Substitution_Rescue'] = (
  1910. ud['sequence'].map(lambda s: _nterm_sub_map.get(s, {}).get('has_rescue', False))
  1911. if 'sequence' in ud.columns
  1912. else pd.Series(False, index=ud.index)
  1913. )
  1914. ud['Nterm_Substitution_Detail'] = (
  1915. ud['sequence'].map(lambda s: _nterm_sub_map.get(s, {}).get('detail', []))
  1916. if 'sequence' in ud.columns
  1917. else pd.Series([[]] * len(ud), index=ud.index)
  1918. )
  1919. # Precomputed here, inside the cached loader, so the N-terminus filter is a
  1920. # column lookup rather than a per-rerun re-parse of the position lists once
  1921. # per cross-filtered facet. See _has_non_nterm_peptide for the rule.
  1922. ud['Has_Non_Nterm_Peptide'] = (
  1923. ud['Tryptic_Start_Positions'].map(_has_non_nterm_peptide)
  1924. if 'Tryptic_Start_Positions' in ud.columns
  1925. else pd.Series(False, index=ud.index)
  1926. )
  1927. ud['Tryptic_Peptide_Count'] = (
  1928. ud['Tryptic_Peptides'].map(_count_peptides)
  1929. if 'Tryptic_Peptides' in ud.columns
  1930. else pd.Series(0, index=ud.index)
  1931. )
  1932. def _sum_semicolon(val):
  1933. if val is None or (isinstance(val, float) and pd.isna(val)):
  1934. return np.nan
  1935. parts = pd.to_numeric(pd.Series(str(val).split(';')).str.strip(), errors='coerce')
  1936. return parts.sum() if parts.notna().any() else np.nan
  1937. def _max_semicolon(val):
  1938. if val is None or (isinstance(val, float) and pd.isna(val)):
  1939. return np.nan
  1940. parts = pd.to_numeric(pd.Series(str(val).split(';')).str.strip(), errors='coerce')
  1941. return parts.max() if parts.notna().any() else np.nan
  1942. if 'Nt_Acetyl_N_PSMs' in ud.columns:
  1943. ud['Nt_Acetyl_Total_PSMs'] = ud['Nt_Acetyl_N_PSMs'].map(_sum_semicolon)
  1944. if 'Nt_Acetyl_PSM_Fraction' in ud.columns:
  1945. ud['Nt_Acetyl_Max_Fraction'] = ud['Nt_Acetyl_PSM_Fraction'].map(_max_semicolon)
  1946. # --- TMT proteomics stats ---
  1947. for src, dest in [
  1948. ('TMT_log2fc_50pct_missing', 'TMT_log2fc_50pct'),
  1949. ('TMT_t_statistic_50pct_missing', 'TMT_t_statistic_50pct'),
  1950. ('TMT_df_50pct_missing', 'TMT_df_50pct'),
  1951. ('TMT_conf_low_50pct_missing', 'TMT_conf_low_50pct'),
  1952. ('TMT_conf_high_50pct_missing', 'TMT_conf_high_50pct'),
  1953. ('TMT_cohens_d_50pct_missing', 'TMT_cohens_d_50pct'),
  1954. ('TMT_pvalue_50pct_missing', 'TMT_pvalue_50pct'),
  1955. ('TMT_qvalue_50pct_missing', 'TMT_qvalue_50pct'),
  1956. ('TMT_log2fc_0pct_missing', 'TMT_log2fc_0pct'),
  1957. ('TMT_t_statistic_0pct_missing', 'TMT_t_statistic_0pct'),
  1958. ('TMT_df_0pct_missing', 'TMT_df_0pct'),
  1959. ('TMT_conf_low_0pct_missing', 'TMT_conf_low_0pct'),
  1960. ('TMT_conf_high_0pct_missing', 'TMT_conf_high_0pct'),
  1961. ('TMT_cohens_d_0pct_missing', 'TMT_cohens_d_0pct'),
  1962. ('TMT_pvalue_0pct_missing', 'TMT_pvalue_0pct'),
  1963. ('TMT_qvalue_0pct_missing', 'TMT_qvalue_0pct'),
  1964. ('rate_control', 'TMT_rate_control'),
  1965. ('rate_ad', 'TMT_rate_ad'),
  1966. ]:
  1967. cols = [c for c in master_df.columns if src in c]
  1968. if cols:
  1969. ud[dest] = pd.to_numeric(master_df[cols[0]], errors='coerce')
  1970. # --- Short-Read RNA (ROSMAP) stats ---
  1971. for src, dest in [
  1972. ('rosmapRNA_log2FoldChange', 'ROSMAP_log2FC'),
  1973. ('rosmapRNA_padj', 'ROSMAP_padj'),
  1974. ('rosmapRNA_pvalue', 'ROSMAP_pvalue'),
  1975. ('rosmapRNA_baseMean', 'ROSMAP_baseMean'),
  1976. ('ROSMAP_BulkRNAseq_CPM', 'ROSMAP_CPM'),
  1977. ('correlation_mainORF_nonAD_rosmap', 'Corr_MainORF_NonAD'),
  1978. ('correlation_mainORF_AD_rosmap', 'Corr_MainORF_AD'),
  1979. ('rosmap_lrt_additive_p', 'RNA_LRT_Add_P'),
  1980. ('rosmap_lrt_interaction_p', 'RNA_LRT_Int_P'),
  1981. ]:
  1982. cols = [c for c in master_df.columns if src in c]
  1983. if cols:
  1984. ud[dest] = pd.to_numeric(master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  1985. # --- Short-Read RNA (MSBB) stats ---
  1986. for src, dest in [
  1987. ('msbbRNA_log2FoldChange', 'MSBB_log2FC'),
  1988. ('msbbRNA_padj', 'MSBB_padj'),
  1989. ('msbbRNA_pvalue', 'MSBB_pvalue'),
  1990. ('msbbRNA_baseMean', 'MSBB_baseMean'),
  1991. ('MSBB_BulkRNAseq_CPM', 'MSBB_CPM'),
  1992. ('correlation_mainORF_nonAD_msbb', 'Corr_MainORF_NonAD_MSBB'),
  1993. ('correlation_mainORF_AD_msbb', 'Corr_MainORF_AD_MSBB'),
  1994. ]:
  1995. cols = [c for c in master_df.columns if src in c]
  1996. if cols:
  1997. ud[dest] = pd.to_numeric(master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  1998. # --- Long-Read RNA (Nanopore) stats ---
  1999. for src, dest in [
  2000. ('nanopore_log2FoldChange', 'Nanopore_log2FC'),
  2001. ('nanopore_padj', 'Nanopore_padj'),
  2002. ('nanopore_pvalue', 'Nanopore_pvalue'),
  2003. ('nanopore_baseMean', 'Nanopore_baseMean'),
  2004. ]:
  2005. cols = [c for c in master_df.columns if src in c]
  2006. if cols:
  2007. ud[dest] = pd.to_numeric(master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  2008. # --- RP3 / Ribo-Seq columns ---
  2009. for src, dest in [
  2010. ('RP3_Default', 'RP3_Default'),
  2011. ('RP3_MM_Amb', 'RP3_MM_Amb'),
  2012. ('RP3_Amb', 'RP3_Amb'),
  2013. ('RP3_MM', 'RP3_MM'),
  2014. ('RiboCode', 'RiboCode'),
  2015. ]:
  2016. cols = [c for c in master_df.columns if c.endswith(src)]
  2017. if cols:
  2018. ud[dest] = master_df[cols].bfill(axis=1).iloc[:, 0]
  2019. for src in PSITE_FRAME_COLS:
  2020. cols = [c for c in master_df.columns if c.endswith(src)]
  2021. if cols:
  2022. ud[src] = pd.to_numeric(master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  2023. # --- BLASTp homology vs UniProt ---
  2024. cols = [c for c in master_df.columns if 'blastp_uniprot_accession_match' in c]
  2025. if cols:
  2026. ud['BLAST_UniProt_Match'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  2027. for src, dest in [
  2028. ('microprotein_percentage_match', 'BLAST_Pct_Match'),
  2029. ('blastp_alignment_length', 'BLAST_Aln_Length'),
  2030. ('evalue', 'BLAST_Evalue'),
  2031. ('blastp_bit', 'BLAST_Bit'),
  2032. ]:
  2033. cols = [c for c in master_df.columns if src in c]
  2034. if cols:
  2035. ud[dest] = pd.to_numeric(master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  2036. # --- scRNA enrichment ---
  2037. if INCLUDE_SCRNA:
  2038. for src, dest in [
  2039. ('scRNA_Enrichment_logFC', 'scRNA_logFC'),
  2040. ('scRNA_Enrichment_p_adj.glb', 'scRNA_padj'),
  2041. ('scRNA_Enrichment_celltype', 'scRNA_celltype'),
  2042. ('scRNA_Enrichment_cell_type_general', 'scRNA_cell_type_general'),
  2043. ]:
  2044. cols = [c for c in master_df.columns if c == src]
  2045. if cols:
  2046. if 'logFC' in src or 'padj' in src:
  2047. ud[dest] = pd.to_numeric(master_df[cols[0]], errors='coerce')
  2048. else:
  2049. ud[dest] = master_df[cols[0]]
  2050. # --- Additional annotation fields for ID card ---
  2051. cols = [c for c in master_df.columns if 'smorf' in c.lower() and 'coordinates' in c.lower()]
  2052. if cols:
  2053. ud['smORF_Coordinates'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  2054. cols = [c for c in master_df.columns if 'ms' in c.lower() and 'evidence' in c.lower()]
  2055. if cols:
  2056. ud['MS_Evidence_Type'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  2057. cols = [c for c in master_df.columns if 'dda' in c.lower() and 'grade' in c.lower()]
  2058. if cols:
  2059. ud['DDA_Grade'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  2060. # --- Kozak context + non-AUG start codon assessment (Salk smORFs only) ---
  2061. for src, dest in [
  2062. ('kozak_full_kozak_strength', 'Kozak_Strength'),
  2063. ('kozak_downstream_kozak_strength', 'Kozak_Downstream_Strength'),
  2064. ('kozak_kozak_class_canonical', 'Kozak_Class_Canonical'),
  2065. ('kozak_kozak_window', 'Kozak_Window'),
  2066. ('nonATG_annotated_start_codon', 'NonATG_Annotated_Codon'),
  2067. ('nonATG_annotated_start_is_initiator', 'NonATG_Is_Initiator'),
  2068. ('nonATG_has_supported_initiation_site', 'NonATG_Has_Supported_Site'),
  2069. ('nonATG_annotated_context_strength', 'NonATG_Context_Strength'),
  2070. ('nonATG_no_site_reason', 'NonATG_No_Site_Reason'),
  2071. # Stringified-list columns (one entry per candidate site), parsed on
  2072. # demand in _precompute_display_columns / the detail card — same
  2073. # convention as Tryptic_Peptides.
  2074. ('nonATG_site_type', 'NonATG_Site_Type'),
  2075. ('nonATG_initiation_codon_position', 'NonATG_Codon_Position'),
  2076. ('nonATG_initiation_codon', 'NonATG_Codon'),
  2077. ('nonATG_cognate_status', 'NonATG_Cognate_Status'),
  2078. ('nonATG_initiation_tier_name', 'NonATG_Tier_Name'),
  2079. ('nonATG_initiator_aa', 'NonATG_Initiator_AA'),
  2080. ('nonATG_predicted_sequence', 'NonATG_Predicted_Sequence'),
  2081. ]:
  2082. cols = [c for c in master_df.columns if c.endswith(src)]
  2083. if cols:
  2084. ud[dest] = master_df[cols].bfill(axis=1).iloc[:, 0]
  2085. cols = [c for c in master_df.columns if c.endswith('nonATG_n_initiation_sites')]
  2086. if cols:
  2087. ud['NonATG_N_Sites'] = pd.to_numeric(master_df[cols].bfill(axis=1).iloc[:, 0], errors='coerce')
  2088. cols = [c for c in master_df.columns if 'database' in c.lower()]
  2089. if cols:
  2090. ud['Database'] = master_df[cols].bfill(axis=1).iloc[:, 0]
  2091. # Keep all database labels seen across merged sources for robust filtering.
  2092. ud['Database_All'] = master_df[cols].apply(
  2093. lambda row: '|'.join(sorted({
  2094. str(v).strip() for v in row
  2095. if pd.notna(v) and str(v).strip() and str(v).strip().lower() != 'none'
  2096. })),
  2097. axis=1
  2098. )
  2099. # --- UniProt annotation score (1–5) for reviewed/TrEMBL microproteins ---
  2100. # Attach UniProt's curation annotation score ONLY where the microprotein IS a
  2101. # genuine UniProt entry (reviewed Swiss-Prot-MP or TrEMBL); novel Salk ORFs stay
  2102. # blank. The microprotein's own accession = its 100%-identity BLASTp self-hit
  2103. # (BLAST_UniProt_Match with BLAST_Pct_Match == 100), so a homolog hit never leaks
  2104. # a neighbour's score onto a novel ORF.
  2105. ud['UniProt_Annotation_Score'] = np.nan
  2106. if 'BLAST_UniProt_Match' in ud.columns:
  2107. _score_map = load_uniprot_annotation_scores()
  2108. if _score_map:
  2109. _is_reviewed = ud.get('Database', pd.Series('', index=ud.index)).astype(str).str.contains(
  2110. 'swiss-prot', case=False, na=False)
  2111. _is_trembl = ud.get('smORF_Class', pd.Series('', index=ud.index)).astype(str).eq('TrEMBL')
  2112. _pct = pd.to_numeric(ud.get('BLAST_Pct_Match', pd.Series(np.nan, index=ud.index)), errors='coerce')
  2113. _self_acc = ud['BLAST_UniProt_Match'].where((_is_reviewed | _is_trembl) & (_pct == 100))
  2114. # UniProt annotation scores are keyed on canonical accessions, so strip any
  2115. # isoform suffix (e.g. Q5BLP8-2 -> Q5BLP8) before mapping.
  2116. _self_acc = _self_acc.astype(str).str.replace(r'-\d+$', '', regex=True)
  2117. ud['UniProt_Annotation_Score'] = pd.to_numeric(
  2118. _self_acc.map(_score_map), errors='coerce')
  2119. # Significance indicators (kept for sidebar filter + metrics)
  2120. ud['TMT_Significant'] = ud.get('TMT_qvalue_50pct', pd.Series(dtype=float)).fillna(1.0) < 0.2
  2121. ud['TMT_Highly_Significant'] = ud.get('TMT_qvalue_50pct', pd.Series(dtype=float)).fillna(1.0) < 0.05
  2122. ud['ROSMAP_Significant'] = ud.get('ROSMAP_padj', pd.Series(dtype=float)).fillna(1.0) < 0.2
  2123. ud['ROSMAP_Highly_Significant'] = ud.get('ROSMAP_padj', pd.Series(dtype=float)).fillna(1.0) < 0.05
  2124. return ud
  2125. # =============================================================================
  2126. # DISK-PARQUET CACHE FOR MERGED + EXTRACTED DATAFRAME
  2127. # =============================================================================
  2128. # The full merge+extract pipeline filters the master dataset and runs many
  2129. # bfill/regex passes (~3-8 s cold). Caching the final unified_df to a single
  2130. # parquet on disk makes subsequent cold starts <0.5 s. Cache key is the mtime
  2131. # fingerprint of every source file — auto-invalidates when any input changes.
  2132. def _source_csv_paths():
  2133. base = Path(__file__).resolve().parent
  2134. return [
  2135. # Single unified source for every master-derived view (microprotein set,
  2136. # tryptic peptides, PROSIT confidence, and all per-analysis columns).
  2137. # Falls back to the committed .zip when the git-ignored .csv is absent
  2138. # (e.g. on Streamlit Cloud) so the fingerprint still tracks the master.
  2139. _resolve_master_source(),
  2140. # Not contained in the master — read directly.
  2141. base / "scRNA_Enrichment" / "scRNA_Enrichment_summary.csv",
  2142. # UniProt annotation-score export (Entry -> Annotation); lives in Code/data
  2143. # so the parquet fingerprint invalidates when it changes.
  2144. base.parent / "Code" / "data" / "uniprotkb_proteome_UP000005640_2026_07_13.tsv",
  2145. # smORF decay-class assignments (NMD × main-ORF disruption); also in
  2146. # Code/data so the parquet fingerprint invalidates when it changes.
  2147. DECAY_CLASS_TSV,
  2148. ]
  2149. @st.cache_data(show_spinner=False)
  2150. def _load_unified_df_with_disk_cache():
  2151. """Load+merge+extract, caching the final dataframe to parquet on disk."""
  2152. cache_dir = Path(__file__).parent / "_cache"
  2153. cache_dir.mkdir(exist_ok=True)
  2154. parquet_path = cache_dir / "unified_df.parquet"
  2155. fp_path = cache_dir / "unified_df.fingerprint"
  2156. # Include mirror_index size in fingerprint so _Spectra_Quality column
  2157. # invalidates when figure libraries change.
  2158. try:
  2159. mirror_index = build_mirror_plot_index()
  2160. except Exception:
  2161. mirror_index = {}
  2162. mirror_fp = f"mirror:{len(mirror_index)}"
  2163. fp = "|".join(
  2164. f"{p.name}:{p.stat().st_mtime_ns}:{p.stat().st_size}"
  2165. for p in _source_csv_paths() if p.exists()
  2166. ) + "|" + mirror_fp + "|" + f"code:{Path(__file__).stat().st_mtime_ns}"
  2167. if parquet_path.exists() and fp_path.exists():
  2168. try:
  2169. if fp_path.read_text() == fp:
  2170. return pd.read_parquet(parquet_path)
  2171. except Exception:
  2172. pass
  2173. master_df = load_and_merge_all_data()
  2174. unified_df = extract_unified_fields(master_df)
  2175. unified_df = _precompute_display_columns(unified_df, mirror_index)
  2176. try:
  2177. unified_df.to_parquet(parquet_path, index=False)
  2178. fp_path.write_text(fp)
  2179. except Exception:
  2180. pass # parquet failures shouldn't block the app
  2181. return unified_df
  2182. def _precompute_display_columns(ud, mirror_index):
  2183. """Pre-compute heavy per-row display columns once, so row-click reruns
  2184. can simply slice instead of running .apply() across all 6.5k rows."""
  2185. # _Spectra_Quality — derive the best PROSIT tier directly from the master's
  2186. # per-peptide `Confidence` list (Strong > Moderate > Weak > Insufficient).
  2187. # This is the same authoritative source used for the annotation `DDA Grade`.
  2188. # For the few rows without a Confidence list (e.g. some Swiss-Prot entries)
  2189. # fall back to the mirror-plot peptide lookup, then default to 'No PROSIT'.
  2190. if 'Confidence' in ud.columns:
  2191. ud['_Spectra_Quality'] = ud['Confidence'].apply(_best_confidence)
  2192. if 'Tryptic_Peptides' in ud.columns and mirror_index:
  2193. missing_mask = ud['_Spectra_Quality'].isna()
  2194. if missing_mask.any():
  2195. ud.loc[missing_mask, '_Spectra_Quality'] = (
  2196. ud.loc[missing_mask, 'Tryptic_Peptides'].apply(
  2197. lambda x: get_spectra_quality(x, mirror_index)
  2198. )
  2199. )
  2200. ud['_Spectra_Quality'] = ud['_Spectra_Quality'].fillna('No PROSIT')
  2201. elif 'Tryptic_Peptides' in ud.columns:
  2202. ud['_Spectra_Quality'] = ud['Tryptic_Peptides'].apply(
  2203. lambda x: get_spectra_quality(x, mirror_index)
  2204. )
  2205. else:
  2206. ud['_Spectra_Quality'] = 'No PROSIT'
  2207. # _Tryptic_Display: joined readable string from list-literal
  2208. def _fmt_peps(val):
  2209. if val is None or (isinstance(val, float) and pd.isna(val)) or val == '':
  2210. return ''
  2211. try:
  2212. peps = ast.literal_eval(str(val))
  2213. if isinstance(peps, str):
  2214. peps = [peps]
  2215. return ' \u00b7 '.join(str(p).strip() for p in peps)
  2216. except (ValueError, SyntaxError):
  2217. return str(val)
  2218. if 'Tryptic_Peptides' in ud.columns:
  2219. ud['_Tryptic_Display'] = ud['Tryptic_Peptides'].map(_fmt_peps)
  2220. # NonATG_Best_Tier: best (lowest-rank) candidate tier for a non-AUG-annotated
  2221. # microprotein, so it can be sorted/filtered as a scalar in the main table.
  2222. _NONATG_TIER_RANK = {
  2223. 'Well-Established Near-Cognate': 0,
  2224. 'Weak Near-Cognate': 1,
  2225. 'Non-Near-Cognate, Strongest Proteomics Evidence': 2,
  2226. 'Non-Near-Cognate, Single-Instance Proteomics Evidence': 3,
  2227. 'Non-Near-Cognate, Inferred (not directly reported as a TIS)': 4,
  2228. }
  2229. def _best_tier(val):
  2230. if val is None or (isinstance(val, float) and pd.isna(val)):
  2231. return pd.NA
  2232. try:
  2233. tiers = ast.literal_eval(str(val))
  2234. except (ValueError, SyntaxError):
  2235. return pd.NA
  2236. if isinstance(tiers, str):
  2237. tiers = [tiers]
  2238. ranked = sorted((t for t in tiers if t in _NONATG_TIER_RANK), key=_NONATG_TIER_RANK.get)
  2239. return ranked[0] if ranked else pd.NA
  2240. if 'NonATG_Tier_Name' in ud.columns:
  2241. ud['NonATG_Best_Tier'] = ud['NonATG_Tier_Name'].map(_best_tier)
  2242. # _NonATG_Candidates_Display: readable "codon@position" list for the main
  2243. # table's Alt-Initiation view, same joined-string convention as _Tryptic_Display.
  2244. def _fmt_candidates(codon_val, pos_val):
  2245. if codon_val is None or (isinstance(codon_val, float) and pd.isna(codon_val)):
  2246. return ''
  2247. try:
  2248. codons = ast.literal_eval(str(codon_val))
  2249. positions = ast.literal_eval(str(pos_val)) if pd.notna(pos_val) else []
  2250. except (ValueError, SyntaxError):
  2251. return ''
  2252. if isinstance(codons, str):
  2253. codons = [codons]
  2254. if isinstance(positions, (str, int)):
  2255. positions = [positions]
  2256. return ' · '.join(
  2257. f"{c}@{p}" if p is not None else str(c)
  2258. for c, p in zip(codons, positions or [None] * len(codons))
  2259. )
  2260. if 'NonATG_Codon' in ud.columns and 'NonATG_Codon_Position' in ud.columns:
  2261. ud['_NonATG_Candidates_Display'] = [
  2262. _fmt_candidates(c, p) for c, p in zip(ud['NonATG_Codon'], ud['NonATG_Codon_Position'])
  2263. ]
  2264. # _UCSC_HTML: pre-built UCSC link URL
  2265. if 'UCSC_Link' in ud.columns:
  2266. ud['_UCSC_HTML'] = ud['UCSC_Link'].apply(
  2267. lambda v: create_ucsc_link({'CLICK_UCSC': v}, CUSTOM_UCSC_SESSION)
  2268. if pd.notna(v) else None
  2269. )
  2270. # _DB_Emoji: vectorized
  2271. db_col = 'Database_All' if 'Database_All' in ud.columns else ('Database' if 'Database' in ud.columns else None)
  2272. if db_col:
  2273. _swiss = ud[db_col].fillna('').str.lower().str.contains('swiss', na=False)
  2274. ud['_DB_Emoji'] = np.where(_swiss, '\U0001F535 Swiss-Prot', '\U0001F7E0 Unreviewed')
  2275. # _TMT_Tier: best (most-stringent) TMT significance tier reached per row.
  2276. # Tier 1: q(50%) < 0.05 · Tier 2: q(50%) < 0.2
  2277. # Tier 3: q(0%) < 0.05 · Tier 4: q(0%) < 0.2
  2278. q50 = pd.to_numeric(ud.get('TMT_qvalue_50pct', pd.Series(index=ud.index, dtype=float)), errors='coerce')
  2279. q0 = pd.to_numeric(ud.get('TMT_qvalue_0pct', pd.Series(index=ud.index, dtype=float)), errors='coerce')
  2280. tmt_tier = pd.Series(pd.NA, index=ud.index, dtype='object')
  2281. tmt_tier[q0 < 0.20] = 'Tier 4'
  2282. tmt_tier[q0 < 0.05] = 'Tier 3'
  2283. tmt_tier[q50 < 0.20] = 'Tier 2'
  2284. tmt_tier[q50 < 0.05] = 'Tier 1'
  2285. ud['_TMT_Tier'] = tmt_tier
  2286. # _RNA_Sig: ROSMAP RNA-seq significance (Significant = padj<0.05, Exploratory = padj<0.2).
  2287. padj = pd.to_numeric(ud.get('ROSMAP_padj', pd.Series(index=ud.index, dtype=float)), errors='coerce')
  2288. rna_sig = pd.Series(pd.NA, index=ud.index, dtype='object')
  2289. rna_sig[padj < 0.20] = 'Exploratory'
  2290. rna_sig[padj < 0.05] = 'Significant'
  2291. ud['_RNA_Sig'] = rna_sig
  2292. # _UniProt_Annotation_Display: gold-star rendering of the 1–5 score.
  2293. if 'UniProt_Annotation_Score' in ud.columns:
  2294. ud['_UniProt_Annotation_Display'] = ud['UniProt_Annotation_Score'].map(_annotation_stars)
  2295. return ud
  2296. # =============================================================================
  2297. # ACTIVE-FILTER SUMMARY
  2298. # =============================================================================
  2299. def _render_active_filter_bar(active_chips, result_count, total_count):
  2300. """Render a compact bar listing the filters currently narrowing the view.
  2301. ``active_chips`` is a list of (label, value) tuples for every sidebar/preset
  2302. control that is active. (The baseline-criteria explainer that used to live
  2303. here as "What am I looking at?" moved to _render_results_table as the
  2304. "Explain the column/variables to me" expander, next to the column-view tabs.)
  2305. """
  2306. if active_chips:
  2307. chips_html = "".join(
  2308. "<span class='filter-chip'>"
  2309. f"<span class='filter-chip-key'>{html.escape(str(k))}</span>"
  2310. f"<span class='filter-chip-val'>{html.escape(str(v))}</span></span>"
  2311. for k, v in active_chips
  2312. )
  2313. st.markdown(
  2314. "<div class='filter-bar'>"
  2315. f"<span class='filter-bar-title'>Active filters ({len(active_chips)})</span>"
  2316. f"{chips_html}"
  2317. "<span class='filter-bar-title' style='margin-left:auto; margin-right:0; "
  2318. f"color:rgba(255,255,255,0.5);'>{result_count:,} of {total_count:,} shown</span>"
  2319. "</div>",
  2320. unsafe_allow_html=True,
  2321. )
  2322. else:
  2323. st.markdown(
  2324. "<div class='filter-bar'><span class='filter-bar-title'>Active filters</span>"
  2325. f"<span class='filter-bar-none'>None \u2014 showing all {total_count:,} "
  2326. "microproteins in the baseline set. Use the Quick-start presets or the sidebar "
  2327. "to refine.</span></div>",
  2328. unsafe_allow_html=True,
  2329. )
  2330. st.caption(
  2331. f"Every view starts from a fixed gold-standard baseline of {total_count:,} brain "
  2332. "microproteins (reviewed Swiss-Prot + unreviewed Salk/TrEMBL), applied before any "
  2333. "sidebar filters — see Figure 1 of the manuscript for the evidence criteria."
  2334. )
  2335. # =============================================================================
  2336. # SIDEBAR FACET HELPERS (UniProt-style checkbox facets with live counts)
  2337. # =============================================================================
  2338. def _is_swiss_mask(df):
  2339. """Boolean mask: True for reviewed Swiss-Prot rows. Single source of truth
  2340. for the Database_All-else-Database fallback, previously duplicated with
  2341. slightly different fallback behavior across several call sites."""
  2342. col = 'Database_All' if 'Database_All' in df.columns else ('Database' if 'Database' in df.columns else None)
  2343. if col is None:
  2344. return pd.Series(False, index=df.index)
  2345. return df[col].fillna('').astype(str).str.lower().str.contains('swiss', na=False)
  2346. # N-terminus filter modes. A tryptic peptide starting at aa 1 or 2 is what Met
  2347. # excision of the *parent* ORF would produce, so on its own it is not evidence
  2348. # that the microprotein is real. These options select the microproteins whose
  2349. # evidence survives discounting that artefact.
  2350. NTERM_MODE_NON_NTERM = "Non-N-terminal peptides only"
  2351. NTERM_MODE_ACETYL_OR_NON_NTERM = "Nt-acetylated, substitution-distinct, or non-N-terminal"
  2352. # Facet domain. Nothing ticked means no filtering; ticked boxes are OR'd, like
  2353. # every other checkbox facet. NON_NTERM is a strict subset of
  2354. # ACETYL_OR_NON_NTERM, so ticking both is the same as ticking the second —
  2355. # exactly how the nested significance tiers (FDR<0.05 within FDR<0.2) behave.
  2356. NTERM_MODES = [NTERM_MODE_NON_NTERM, NTERM_MODE_ACETYL_OR_NON_NTERM]
  2357. def _count_checkbox(key, label, n, help=None):
  2358. """A facet checkbox whose live count lives in the label, not the key.
  2359. The `value=` seeding is load-bearing and the same trick as in
  2360. _render_facet_checkboxes: Streamlit derives widget identity partly from the
  2361. label, so a count baked into it mints a new widget on every recount, and a
  2362. new widget falls back to its default — wiping the tick even though `key` is
  2363. stable. Seeding the default from stored state makes that fallback a no-op.
  2364. """
  2365. st.checkbox(
  2366. f"{label} ({n:,})", key=key,
  2367. value=bool(st.session_state.get(key, False)),
  2368. help=help,
  2369. )
  2370. def _nterm_mask(df, mode):
  2371. """Row mask for an N-terminus mode; None means no filtering.
  2372. NON_NTERM keeps rows with a tryptic peptide starting at aa >= 3 — evidence
  2373. beyond the Met-excision artefact. ACETYL_OR_NON_NTERM also readmits two
  2374. other kinds of rows: Nt-acetylated (Nt-acetylation is co-translational and
  2375. marks a genuine protein N-terminus, so an acetylated aa-1/2 peptide is real
  2376. evidence), and rows whose N-terminal peptide carries 1-2 amino-acid
  2377. substitutions vs. its matched UniProt isoform (Nterm_Substitution_Rescue,
  2378. see compute_nterm_peptide_substitutions.py) — a peptide that differs from
  2379. the canonical protein's sequence at that position cannot be a bare
  2380. Met-excision fragment of it.
  2381. Rows with no tryptic peptides at all fail both modes — absence of peptide
  2382. data is not evidence.
  2383. """
  2384. if mode not in (NTERM_MODE_NON_NTERM, NTERM_MODE_ACETYL_OR_NON_NTERM):
  2385. return None
  2386. mask = _non_nterm_series(df)
  2387. if mask is None:
  2388. return None # no peptide data at all in this frame
  2389. if mode == NTERM_MODE_ACETYL_OR_NON_NTERM:
  2390. if 'Nt_Acetylated' in df.columns:
  2391. mask = mask | df['Nt_Acetylated'].fillna(False).astype(bool)
  2392. if 'Nterm_Substitution_Rescue' in df.columns:
  2393. mask = mask | df['Nterm_Substitution_Rescue'].fillna(False).astype(bool)
  2394. return mask
  2395. def _render_facet_checkboxes(series, key_prefix, order=None, help=None, label_fn=None,
  2396. domain=None):
  2397. """Render one st.checkbox per distinct value in `series`, each labeled with
  2398. a live count. Sorted by count descending unless an explicit `order` list is
  2399. given (for facets with a real ordinal/curated meaning, e.g. quality tiers).
  2400. `series` is the CROSS-FILTERED view for this facet — the data narrowed by
  2401. every OTHER active filter but not by this facet's own selection. That is
  2402. what makes these counts react to boxes checked anywhere else (including in
  2403. facets rendered further down the sidebar), while a facet's own boxes never
  2404. zero out their own counts.
  2405. Returns the list of currently-checked values — the same list shape
  2406. st.multiselect used to return, so downstream `.isin(selected)` filtering
  2407. code needs no changes, only the widget call site does.
  2408. `label_fn`, if given, formats each value for display only (e.g. to
  2409. capitalize "strong" -> "Strong"); the underlying value used for the
  2410. checkbox key and the returned `selected` list is unaffected.
  2411. `domain` (or `order`) is the facet's full value list, and every value in it
  2412. gets a box even at count 0. A facet must never drop an option as other
  2413. filters narrow the data: a vanished box cannot be ticked, so the only way
  2414. back to it would be to clear the filters that hid it. Zero-count boxes stay
  2415. clickable and simply read "(0)".
  2416. """
  2417. counts = series.value_counts(dropna=True)
  2418. # Curated `order` is rendered whole; otherwise sort by count descending and
  2419. # append any declared domain value the current view happens not to contain.
  2420. values = list(order) if order else counts.index.tolist()
  2421. for v in (domain or []):
  2422. if v not in values:
  2423. values.append(v)
  2424. selected = []
  2425. for i, v in enumerate(values):
  2426. n = int(counts.get(v, 0))
  2427. label = label_fn(v) if label_fn else v
  2428. _key = f"{key_prefix}_{v}"
  2429. # `value=` is seeded from the box's own stored state, and that is load-
  2430. # bearing: Streamlit derives widget identity partly from the LABEL, so
  2431. # the live count baked into it mints a new widget every time any other
  2432. # facet changes the count — and a new widget silently falls back to its
  2433. # default, wiping the tick even though `key` is stable. Seeding the
  2434. # default with the stored value makes that fallback a no-op. Without
  2435. # this, each click clears every other facet's selections.
  2436. checked = st.checkbox(
  2437. f"{label} ({n:,})", key=_key,
  2438. value=bool(st.session_state.get(_key, False)),
  2439. help=(help if i == 0 else None),
  2440. )
  2441. if checked:
  2442. selected.append(v)
  2443. return selected
  2444. # =============================================================================
  2445. # MAIN APP — SINGLE-PAGE LAYOUT
  2446. # =============================================================================
  2447. # ── Filter-state persistence across the entry page ──────────────────────────
  2448. # Streamlit drops widget state for any widget that did not render on a rerun,
  2449. # and the entry page returns from main() before the sidebar draws. Without a
  2450. # shadow copy, opening a microprotein silently clears every filter and the
  2451. # search box, so clicking "Back to results" lands on an unfiltered table —
  2452. # measured: a checked facet reads True during the entry render and False on the
  2453. # way back. The shadow key is a plain (non-widget) key, so nothing culls it.
  2454. #
  2455. # Not recoverable this way: the table's column sort, which st.dataframe never
  2456. # reports back to Python, and scroll position. Both are lost by design.
  2457. _PERSIST_SHADOW = '_filter_state_shadow'
  2458. _PERSIST_EXTRA = ('hero_search',)
  2459. def _persisted_filter_keys():
  2460. """Widget keys whose values should survive an entry-page detour."""
  2461. return [
  2462. k for k in st.session_state
  2463. if isinstance(k, str) and (k.startswith('f_') or k in _PERSIST_EXTRA)
  2464. ]
  2465. def _snapshot_filter_state():
  2466. """Mirror live filter widgets into the shadow key. Call after they render."""
  2467. st.session_state[_PERSIST_SHADOW] = {
  2468. k: st.session_state[k] for k in _persisted_filter_keys()
  2469. }
  2470. def _restore_filter_state():
  2471. """Re-seed filter widgets from the shadow key.
  2472. Must run *before* the widgets instantiate — assigning to a widget key after
  2473. its widget exists on the same run raises. Only fills keys Streamlit culled,
  2474. so it never clobbers a selection the user is actively changing.
  2475. """
  2476. saved = st.session_state.get(_PERSIST_SHADOW)
  2477. if not saved:
  2478. return
  2479. for k, v in saved.items():
  2480. if k not in st.session_state:
  2481. st.session_state[k] = v
  2482. def main():
  2483. # ── Gradient banner placeholder (rendered after data loads) ──
  2484. header_slot = st.empty()
  2485. # Load data (parquet-cached on disk; auto-rebuilds when source CSVs change)
  2486. with st.spinner("Loading microprotein database..."):
  2487. try:
  2488. unified_df = _load_unified_df_with_disk_cache()
  2489. if 'Annotation_Status' in unified_df.columns:
  2490. # Keep Swiss-Prot rows (which have no Annotation_Status) alongside annotated rows.
  2491. # Only include Swiss-Prot entries that have proteomics evidence (i.e., appear in
  2492. # the Proteomics CSV). The 752 Swiss-Prot-MP entries added after first submission
  2493. # have no MS data and should not appear in the default view.
  2494. has_annotation = ~(unified_df['Annotation_Status'].isna() | (unified_df['Annotation_Status'] == 'None'))
  2495. is_swiss = unified_df.get('Database_All', unified_df.get('Database', pd.Series('', index=unified_df.index))).fillna('').str.lower().str.contains('swiss')
  2496. _prot_present_col = next((c for c in unified_df.columns if 'proteomics' in c.lower() and c.endswith('_present')), None)
  2497. if _prot_present_col:
  2498. has_proteomics = unified_df[_prot_present_col].fillna(False).astype(bool)
  2499. is_swiss = is_swiss & has_proteomics
  2500. unified_df = unified_df[has_annotation | is_swiss]
  2501. if unified_df.empty:
  2502. st.error("No data could be loaded.")
  2503. return
  2504. except Exception as e:
  2505. st.error(f"Error loading data: {e}")
  2506. return
  2507. # ── Pre-computed indexes (mirror plots, expression profiles, coords) ──
  2508. mirror_index = build_mirror_plot_index()
  2509. expression_index = build_expression_profile_index()
  2510. seq_to_coords = build_seq_to_coords_index()
  2511. # _Spectra_Quality is now precomputed inside the cached load (see
  2512. # _precompute_display_columns); fall back only if the column is missing.
  2513. if '_Spectra_Quality' not in unified_df.columns:
  2514. if 'Tryptic_Peptides' in unified_df.columns and mirror_index:
  2515. unified_df['_Spectra_Quality'] = unified_df['Tryptic_Peptides'].apply(
  2516. lambda x: get_spectra_quality(x, mirror_index)
  2517. )
  2518. else:
  2519. unified_df['_Spectra_Quality'] = unified_df.get(
  2520. 'Tryptic_Peptides', pd.Series(dtype=object)
  2521. ).apply(lambda x: 'Insufficient' if (x and not pd.isna(x)) else 'No PROSIT')
  2522. # ── Route: ?mp=<entry id> takes over the whole page with that microprotein's
  2523. # entry, replacing the table and the filter sidebar. Checked before any
  2524. # of them render, so the entry page skips the cross-filter engine and the
  2525. # display-table prep entirely. Resolved against the *unfiltered* frame so
  2526. # a shared link opens regardless of the recipient's filter state. ──
  2527. _entry_key = st.query_params.get(ENTRY_QUERY_PARAM)
  2528. if _entry_key:
  2529. _render_entry_page(unified_df, _entry_key, mirror_index,
  2530. expression_index, seq_to_coords)
  2531. return
  2532. # Back on the table path: re-seed anything the entry-page detour culled.
  2533. # Must precede the search box and the sidebar, which are the widgets it
  2534. # restores.
  2535. _restore_filter_state()
  2536. # ── Search-first landing strip (a distinct "sandbox" zone to explore by example) ──
  2537. # NOTE: filter session-state keys all use the "f_" prefix by convention.
  2538. st.session_state.setdefault("quick_preset", None)
  2539. st.markdown(
  2540. "<style>div[data-testid='stVerticalBlockBorderWrapper']:has(.hero-sandbox-marker){"
  2541. "background:linear-gradient(135deg,rgba(116,162,183,0.16),rgba(237,134,81,0.07))!important;"
  2542. "border:1.5px dashed rgba(116,162,183,0.7)!important;"
  2543. "border-radius:14px!important;padding:0.8rem 1.15rem 0.55rem!important;"
  2544. "box-shadow:0 2px 16px rgba(0,0,0,0.18)!important;margin-bottom:0.8rem!important;}</style>",
  2545. unsafe_allow_html=True,
  2546. )
  2547. with st.container(border=True):
  2548. st.markdown(
  2549. "<div class='hero-sandbox-marker' style='font-size:0.78rem; font-weight:700; "
  2550. "letter-spacing:0.05em; text-transform:uppercase; color:#8da8b8; margin-bottom:0.25rem;'>"
  2551. "Quick start \u00b7 example views</div>",
  2552. unsafe_allow_html=True,
  2553. )
  2554. hero_search = st.text_input(
  2555. "Search microproteins",
  2556. key="hero_search",
  2557. placeholder="Search a gene (e.g. BCL3) or paste an amino-acid sequence (e.g. MAASGK)\u2026",
  2558. label_visibility="collapsed",
  2559. )
  2560. st.caption(
  2561. f"Explore {len(unified_df):,} brain microproteins. These quick views are just examples to "
  2562. "toy with \u2014 or build your own with the search above and the sidebar filters."
  2563. )
  2564. _quick_presets = [
  2565. ("lncRNA Microproteins", "lncrna"),
  2566. ("Strong Graded Microproteins", "strong_graded"),
  2567. ]
  2568. _active_preset = st.session_state.get("quick_preset")
  2569. _preset_cols = st.columns(len(_quick_presets) + 1)
  2570. for _pcol, (_plabel, _pkey) in zip(_preset_cols, _quick_presets):
  2571. if _pcol.button(
  2572. _plabel,
  2573. key=f"qp_{_pkey}",
  2574. use_container_width=True,
  2575. type=("primary" if _active_preset == _pkey else "secondary"),
  2576. ):
  2577. st.session_state["quick_preset"] = None if _active_preset == _pkey else _pkey
  2578. st.rerun()
  2579. if _preset_cols[-1].button(
  2580. "Show all",
  2581. key="qp_clear",
  2582. use_container_width=True,
  2583. disabled=not _active_preset,
  2584. ):
  2585. st.session_state["quick_preset"] = None
  2586. st.rerun()
  2587. # ── Sidebar: ALL filters (UniProt-style checkbox facets with live counts).
  2588. # Filtering happens INLINE here, narrowing a running `filtered_df`
  2589. # sequentially as each section renders, so every facet's counts
  2590. # reflect everything selected in the sections above it. ──
  2591. with st.sidebar:
  2592. st.markdown('<div style="font-size:1.05rem; font-weight:700; color:#ffffff; margin-bottom:0.7rem; padding-bottom:0.4rem; border-bottom:1px solid rgba(116,162,183,0.25);">Filters</div>', unsafe_allow_html=True)
  2593. st.caption("Search is at the top of the page. Use these controls to refine results.")
  2594. base_df = unified_df.copy()
  2595. # Unified search box (top of page): matches parent gene OR amino-acid sequence.
  2596. if hero_search:
  2597. _sq = hero_search.strip()
  2598. _smask = pd.Series(False, index=base_df.index)
  2599. if 'Parent_Gene' in base_df.columns:
  2600. _smask = _smask | base_df['Parent_Gene'].str.contains(_sq, case=False, na=False)
  2601. if 'sequence' in base_df.columns:
  2602. _smask = _smask | base_df['sequence'].str.contains(_sq, case=False, na=False)
  2603. base_df = base_df[_smask]
  2604. # Curated quick-view preset (landing-strip buttons) — applied first so
  2605. # it acts as a coarse pre-filter and every facet below shows counts
  2606. # consistent with an active preset.
  2607. _qp = st.session_state.get("quick_preset")
  2608. if _qp == "lncrna":
  2609. if 'smORF_Group' in base_df.columns:
  2610. base_df = base_df[base_df['smORF_Group'] == 'lncRNA']
  2611. elif _qp == "strong_graded":
  2612. if '_Spectra_Quality' in base_df.columns:
  2613. base_df = base_df[base_df['_Spectra_Quality'] == 'Strong']
  2614. # ── Cross-filtering engine ───────────────────────────────────────────
  2615. # Every filter is turned into a row mask over `base_df` BEFORE any widget
  2616. # renders (selections are read straight from session_state, whose keys
  2617. # Streamlit has already updated for this rerun). A facet is then drawn
  2618. # against `_narrow(skip=<itself>)` — the data narrowed by all the *other*
  2619. # filters — so checking a box updates the counts in every other facet,
  2620. # including the ones above it, instead of only the ones below.
  2621. _all_true = pd.Series(True, index=base_df.index)
  2622. def _ss(key, default=None):
  2623. return st.session_state.get(key, default)
  2624. def _sel(prefix, domain):
  2625. """Values in `domain` whose checkbox is currently ticked."""
  2626. return [v for v in domain if st.session_state.get(f"{prefix}_{v}", False)]
  2627. def _col(name):
  2628. return base_df[name] if name in base_df.columns else None
  2629. # Stable value domains (independent of what is currently filtered).
  2630. _dom_status = ['Reviewed (Swiss-Prot)', 'Unreviewed (Salk/TrEMBL)']
  2631. _dom_group = list(SMORF_PARENT_GROUPS.keys())
  2632. _dom_dsub = list(SMORF_PARENT_GROUPS['Downstream'])
  2633. _dom_codon = ['ATG', 'nonATG']
  2634. _dom_kozak = ['strong', 'adequate', 'weak']
  2635. _dom_quality = list(QUALITY_LEVELS)
  2636. _dom_decay = list(DECAY_CLASS_LEVELS)
  2637. _dom_ribo = ['RiboCode-SAM', 'Coverage']
  2638. _dom_shortstop = (sorted(unified_df['ShortStop_Label'].dropna().unique())
  2639. if 'ShortStop_Label' in unified_df.columns else [])
  2640. _dom_tmt_sig = ['Tier 1', 'Tier 2', 'Tier 3', 'Tier 4']
  2641. _dom_rna_sig = ['Significant', 'Exploratory']
  2642. selected_status = _sel('f_status', _dom_status)
  2643. selected_groups = _sel('f_smorf_group', _dom_group)
  2644. # Sub-type only exists under Downstream; ignore stale ticks otherwise.
  2645. selected_downstream_sub = (_sel('f_downstream_sub', _dom_dsub)
  2646. if 'Downstream' in selected_groups else [])
  2647. selected_codons = _sel('f_codons', _dom_codon)
  2648. selected_kozak = _sel('f_kozak', _dom_kozak)
  2649. selected_ribo = _sel('f_ribo', _dom_ribo)
  2650. selected_quality = _sel('f_quality', _dom_quality)
  2651. selected_min_peptides = _sel('f_min_peptides', MIN_PEPTIDE_TIERS)
  2652. selected_decay = _sel('f_decay', _dom_decay)
  2653. selected_shortstop = _sel('f_shortstop', _dom_shortstop)
  2654. selected_tmt_sig = _sel('f_tmt_sig', _dom_tmt_sig)
  2655. selected_rna_sig = _sel('f_rna_sig', _dom_rna_sig)
  2656. # Slider bounds come from the full dataset so they never move underfoot.
  2657. def _bounds(col, cast):
  2658. if col not in unified_df.columns:
  2659. return None
  2660. data = unified_df[col].dropna()
  2661. if data.empty:
  2662. return None
  2663. return cast(data.min()), cast(data.max())
  2664. _len_bounds = _bounds('Protein_Length', int)
  2665. _spec_bounds = _bounds('Unique_Spectral_Counts', int)
  2666. _phylo_bounds = _bounds('PhyloCSF_Score', float)
  2667. _uni_bounds = ((1, 5) if ('UniProt_Annotation_Score' in unified_df.columns
  2668. and unified_df['UniProt_Annotation_Score'].notna().any())
  2669. else None)
  2670. length_range = _ss('f_length_range', _len_bounds) if _len_bounds else None
  2671. spectral_range = _ss('f_spectral_range', _spec_bounds) if _spec_bounds else None
  2672. phylocsf_range = _ss('f_phylocsf_range', _phylo_bounds) if _phylo_bounds else None
  2673. uniprot_score_range = _ss('f_uniprot_range', _uni_bounds) if _uni_bounds else None
  2674. selected_nterm = _sel('f_nterm', NTERM_MODES)
  2675. _masks = {}
  2676. def _add_mask(name, mask):
  2677. if mask is not None:
  2678. _masks[name] = mask
  2679. # Status — one box ticked narrows; none or both means "show all".
  2680. if len(selected_status) == 1:
  2681. _sw = _is_swiss_mask(base_df)
  2682. _add_mask('status', _sw if selected_status[0].startswith('Reviewed') else ~_sw)
  2683. # smORF group + nested Downstream sub-type.
  2684. if selected_groups and 'smORF_Class' in base_df.columns:
  2685. _allowed = []
  2686. for _g in selected_groups:
  2687. _allowed.extend(SMORF_PARENT_GROUPS.get(_g, []))
  2688. _add_mask('group', base_df['smORF_Class'].isin(_allowed))
  2689. if selected_downstream_sub and 'smORF_Class' in base_df.columns:
  2690. _dall = set(SMORF_PARENT_GROUPS['Downstream'])
  2691. _add_mask('dsub', (~base_df['smORF_Class'].isin(_dall))
  2692. | base_df['smORF_Class'].isin(selected_downstream_sub))
  2693. # Start codon / Kozak — rows with no value are always retained.
  2694. if selected_codons and _col('Start_Codon') is not None:
  2695. _add_mask('codon', base_df['Start_Codon'].isin(selected_codons)
  2696. | base_df['Start_Codon'].isna())
  2697. if selected_kozak and _col('Kozak_Strength') is not None:
  2698. _add_mask('kozak', base_df['Kozak_Strength'].isin(selected_kozak)
  2699. | base_df['Kozak_Strength'].isna())
  2700. # Ribosome coverage: RiboCode-SAM = RiboCode ORF call with a ShortStop
  2701. # SAM label; Coverage = P-sites in at least one frame (% frame 0, 1 or 2 > 0).
  2702. def _ribo_mask(df, lvl):
  2703. if lvl == 'RiboCode-SAM':
  2704. if 'RiboCode' not in df.columns or 'ShortStop_Label' not in df.columns:
  2705. return None
  2706. return ((df['RiboCode'].astype(str).str.strip().str.upper() == 'TRUE')
  2707. & df['ShortStop_Label'].astype(str).str.startswith('SAM'))
  2708. _pct = [f'Psites_pct_frame{i}' for i in range(3)]
  2709. if not all(c in df.columns for c in _pct):
  2710. return None
  2711. return (df[_pct].apply(pd.to_numeric, errors='coerce') > 0).any(axis=1)
  2712. _ribo_any = None
  2713. for _lvl in selected_ribo:
  2714. _m = _ribo_mask(base_df, _lvl)
  2715. if _m is None:
  2716. continue
  2717. _ribo_any = _m if _ribo_any is None else (_ribo_any | _m)
  2718. _add_mask('ribo', _ribo_any)
  2719. if selected_quality and _col('_Spectra_Quality') is not None:
  2720. _add_mask('quality', base_df['_Spectra_Quality'].isin(selected_quality))
  2721. if selected_min_peptides:
  2722. _pc = _peptide_count_series(base_df)
  2723. if _pc is not None:
  2724. _add_mask('peptides', _pc >= min(selected_min_peptides))
  2725. if selected_decay and _col('NMD_Decay_Class') is not None:
  2726. _add_mask('decay', base_df['NMD_Decay_Class'].isin(selected_decay))
  2727. if selected_shortstop and _col('ShortStop_Label') is not None:
  2728. _add_mask('shortstop', base_df['ShortStop_Label'].isin(selected_shortstop))
  2729. # Significance facets. Tier 1 = strongest (q<0.05, ≥50% samples) … Tier 4
  2730. # = weakest (q<0.2, ≥1 sample). Checked boxes are OR'd, matching every
  2731. # other facet in the sidebar; no box ticked means no significance filter.
  2732. # Matching on the stable tier token means a changed emoji or count in the
  2733. # label can never silently disable a filter.
  2734. _tmt_tiers = {
  2735. 'Tier 1': ('TMT_qvalue_50pct', 0.05),
  2736. 'Tier 2': ('TMT_qvalue_50pct', 0.2),
  2737. 'Tier 3': ('TMT_qvalue_0pct', 0.05),
  2738. 'Tier 4': ('TMT_qvalue_0pct', 0.2),
  2739. }
  2740. _tmt_any = None
  2741. for _tier in selected_tmt_sig:
  2742. _qcol, _thr = _tmt_tiers[_tier]
  2743. if _qcol not in base_df.columns:
  2744. continue
  2745. _m = pd.to_numeric(base_df[_qcol], errors='coerce') < _thr
  2746. _tmt_any = _m if _tmt_any is None else (_tmt_any | _m)
  2747. _add_mask('tmt', _tmt_any)
  2748. _rna_cols = {
  2749. 'Significant': 'ROSMAP_Highly_Significant',
  2750. 'Exploratory': 'ROSMAP_Significant',
  2751. }
  2752. _rna_any = None
  2753. for _lvl in selected_rna_sig:
  2754. _rcol = _rna_cols[_lvl]
  2755. if _rcol not in base_df.columns:
  2756. continue
  2757. _m = base_df[_rcol].fillna(False).astype(bool)
  2758. _rna_any = _m if _rna_any is None else (_rna_any | _m)
  2759. _add_mask('rna', _rna_any)
  2760. # Numeric ranges — only a mask when the slider actually narrows, and rows
  2761. # with no value are always retained.
  2762. def _range_mask(name, col, rng, bounds):
  2763. if not rng or not bounds or col not in base_df.columns:
  2764. return
  2765. if rng[0] <= bounds[0] and rng[1] >= bounds[1]:
  2766. return
  2767. _add_mask(name, base_df[col].between(rng[0], rng[1]) | base_df[col].isna())
  2768. _range_mask('length', 'Protein_Length', length_range, _len_bounds)
  2769. _range_mask('spectral', 'Unique_Spectral_Counts', spectral_range, _spec_bounds)
  2770. _range_mask('phylocsf', 'PhyloCSF_Score', phylocsf_range, _phylo_bounds)
  2771. _range_mask('uniprot', 'UniProt_Annotation_Score', uniprot_score_range, _uni_bounds)
  2772. # N-terminus: discount the Met-excision artefact of the parent ORF.
  2773. # _add_mask ignores None, so "All microproteins" needs no branch here.
  2774. _nt_any = None
  2775. for _mode in selected_nterm:
  2776. _m = _nterm_mask(base_df, _mode)
  2777. if _m is None:
  2778. continue
  2779. _nt_any = _m if _nt_any is None else (_nt_any | _m)
  2780. _add_mask('nterm', _nt_any)
  2781. def _narrow(skip=None):
  2782. """base_df narrowed by every active filter except the named one(s)."""
  2783. skip = {skip} if isinstance(skip, str) else set(skip or ())
  2784. m = _all_true
  2785. for _name, _mk in _masks.items():
  2786. if _name not in skip:
  2787. m = m & _mk
  2788. return base_df[m]
  2789. # The fully filtered result — every widget below is drawn against a
  2790. # cross-filtered view, so this is computed up front rather than
  2791. # accumulated as the sidebar renders.
  2792. filtered_df = _narrow()
  2793. # ── Status ──
  2794. if 'Database_All' in base_df.columns or 'Database' in base_df.columns:
  2795. with st.expander("Status", expanded=False):
  2796. _v = _narrow(skip='status')
  2797. _status_series = pd.Series(
  2798. np.where(_is_swiss_mask(_v), 'Reviewed (Swiss-Prot)', 'Unreviewed (Salk/TrEMBL)'),
  2799. index=_v.index,
  2800. )
  2801. _render_facet_checkboxes(
  2802. _status_series, key_prefix='f_status',
  2803. order=_dom_status,
  2804. help="Check one to narrow; check both (or neither) to show all.",
  2805. )
  2806. # ── smORF Type (+ nested Downstream Sub-type) ──
  2807. if 'smORF_Group' in base_df.columns:
  2808. with st.expander("smORF Type", expanded=False):
  2809. _render_facet_checkboxes(
  2810. _narrow(skip={'group', 'dsub'})['smORF_Group'], key_prefix='f_smorf_group',
  2811. order=_dom_group,
  2812. help="Top-level smORF category.",
  2813. )
  2814. if 'Downstream' in selected_groups and 'smORF_Class' in base_df.columns:
  2815. _dv = _narrow(skip='dsub')
  2816. _dv = _dv[_dv['smORF_Class'].isin(_dom_dsub)]
  2817. if _dv['smORF_Class'].notna().any() or selected_downstream_sub:
  2818. st.markdown("**Downstream Sub-type**")
  2819. _render_facet_checkboxes(
  2820. _dv['smORF_Class'], key_prefix='f_downstream_sub',
  2821. order=_dom_dsub,
  2822. help="Refine within Downstream ORFs (doORF, daORF, etc.)",
  2823. )
  2824. # ── Evidence & Quality ──
  2825. with st.expander("Evidence & Quality", expanded=False):
  2826. st.markdown("**Spectra Quality**")
  2827. _render_facet_checkboxes(
  2828. _narrow(skip='quality')['_Spectra_Quality'], key_prefix='f_quality',
  2829. order=_dom_quality,
  2830. help="Best Prosit spectral-match confidence tier (from the master Confidence column).",
  2831. )
  2832. st.markdown("**Ribosome Coverage**")
  2833. _rv = _narrow(skip='ribo')
  2834. _ribo_help = {
  2835. 'RiboCode-SAM': "RiboCode ORF call with a ShortStop SAM (Secreted/Intracellular) label.",
  2836. 'Coverage': "P-sites in at least one frame (P-site % Frame 0, 1 or 2 > 0).",
  2837. }
  2838. for _lvl in _dom_ribo:
  2839. _m = _ribo_mask(_rv, _lvl)
  2840. _count_checkbox(f"f_ribo_{_lvl}", _lvl,
  2841. int(_m.sum()) if _m is not None else 0,
  2842. help=_ribo_help[_lvl])
  2843. st.markdown("**Peptide Evidence**")
  2844. _pc = _peptide_count_series(_narrow(skip='peptides'))
  2845. for _i, _n in enumerate(MIN_PEPTIDE_TIERS):
  2846. _count_checkbox(f"f_min_peptides_{_n}", f"≥{_n} peptides",
  2847. int((_pc >= _n).sum()) if _pc is not None else 0,
  2848. help=("Keep microproteins identified by at least this many "
  2849. "distinct tryptic peptide sequences (missed-cleavage "
  2850. "variants count as distinct sequences). Tiers are nested, "
  2851. "so ticking several is the same as ticking the lowest."
  2852. if _i == 0 else None))
  2853. # ── Differential Expression ──
  2854. # Checkbox facets like the rest of the sidebar: the checkbox *key* carries
  2855. # the stable tier token and the live count lives only in the label.
  2856. with st.expander("Differential Expression", expanded=False):
  2857. _sig_checkbox = _count_checkbox
  2858. has_tmt_50 = 'TMT_qvalue_50pct' in base_df.columns
  2859. has_tmt_0 = 'TMT_qvalue_0pct' in base_df.columns
  2860. if has_tmt_50 or has_tmt_0:
  2861. _tv = _narrow(skip='tmt')
  2862. _tmt_help = (
  2863. "Filter by TMT proteomics q-value at the chosen stringency; "
  2864. "ticking several tiers keeps microproteins in any of them. "
  2865. "≥50% samples (stringent): peptide quantified in at least half the donors per condition. "
  2866. "≥1 sample/condition: peptide quantified in at least one donor per condition (lenient)."
  2867. )
  2868. st.markdown("**TMT-MS Significance**")
  2869. if has_tmt_50:
  2870. _q50 = pd.to_numeric(_tv['TMT_qvalue_50pct'], errors='coerce')
  2871. _sig_checkbox("f_tmt_sig_Tier 1", "Tier 1 — FDR < 0.05, ≥50% samples",
  2872. int((_q50 < 0.05).sum()), help=_tmt_help)
  2873. _sig_checkbox("f_tmt_sig_Tier 2", "Tier 2 — FDR < 0.2, ≥50% samples",
  2874. int((_q50 < 0.2).sum()))
  2875. if has_tmt_0:
  2876. _q0 = pd.to_numeric(_tv['TMT_qvalue_0pct'], errors='coerce')
  2877. _sig_checkbox("f_tmt_sig_Tier 3", "Tier 3 — FDR < 0.05, ≥1 sample/condition",
  2878. int((_q0 < 0.05).sum()), help=(_tmt_help if not has_tmt_50 else None))
  2879. _sig_checkbox("f_tmt_sig_Tier 4", "Tier 4 — FDR < 0.2, ≥1 sample/condition",
  2880. int((_q0 < 0.2).sum()))
  2881. if 'ROSMAP_Significant' in base_df.columns:
  2882. _rv = _narrow(skip='rna')
  2883. _rna_exp_n = int(_rv['ROSMAP_Significant'].sum())
  2884. _rna_sig_n = int(_rv['ROSMAP_Highly_Significant'].sum()) if 'ROSMAP_Highly_Significant' in _rv.columns else 0
  2885. st.markdown("**RNA Significance**")
  2886. _sig_checkbox("f_rna_sig_Significant", "Significant (FDR < 0.05)", _rna_sig_n,
  2887. help="Filter by ROSMAP bulk RNA-seq adjusted p-value")
  2888. _sig_checkbox("f_rna_sig_Exploratory", "Exploratory (FDR < 0.2)", _rna_exp_n)
  2889. # ── ShortStop Label ──
  2890. if 'ShortStop_Label' in base_df.columns:
  2891. with st.expander("ShortStop Label", expanded=False):
  2892. _render_facet_checkboxes(
  2893. _narrow(skip='shortstop')['ShortStop_Label'], key_prefix='f_shortstop',
  2894. domain=_dom_shortstop,
  2895. help="ShortStop ribosome profiling classification",
  2896. )
  2897. # ── Score & Length Ranges ──
  2898. # Each "in range" caption counts against the view narrowed by every other
  2899. # filter but not by this slider, matching the facet counts above.
  2900. with st.expander("Score & Length Ranges", expanded=False):
  2901. def _range_slider(label, name, col, bounds, step=None, help=None):
  2902. if not bounds or col not in base_df.columns:
  2903. return
  2904. st.slider(label, min_value=bounds[0], max_value=bounds[1],
  2905. value=bounds, step=step, key=f"f_{name}_range", help=help)
  2906. _rv = _narrow(skip=name)
  2907. _rng = _ss(f"f_{name}_range", bounds)
  2908. _n = int((_rv[col].between(_rng[0], _rng[1]) | _rv[col].isna()).sum())
  2909. st.caption(f"→ {_n:,} microproteins in range")
  2910. _range_slider("Protein Length (aa)", "length", 'Protein_Length', _len_bounds)
  2911. _range_slider("Unique Spectral Counts", "spectral", 'Unique_Spectral_Counts', _spec_bounds)
  2912. _range_slider("PhyloCSF Score", "phylocsf", 'PhyloCSF_Score', _phylo_bounds, step=0.5)
  2913. _range_slider(
  2914. "UniProt Annotation Score (★)", "uniprot", 'UniProt_Annotation_Score',
  2915. _uni_bounds, step=1,
  2916. help="UniProt curation annotation score (1–5). Reviewed / TrEMBL "
  2917. "entries only; microproteins without a UniProt score are always retained.",
  2918. )
  2919. # ── Start Codon ──
  2920. if 'Start_Codon' in base_df.columns:
  2921. with st.expander("Start Codon", expanded=False):
  2922. _render_facet_checkboxes(
  2923. _narrow(skip='codon')['Start_Codon'], key_prefix='f_codons',
  2924. order=_dom_codon,
  2925. help="ATG (canonical) vs nonATG (annotated with a near-cognate/non-standard "
  2926. "start — see the Alt-Initiation column view for detail).",
  2927. )
  2928. # ── ORF Rules (NMD × main-ORF disruption) ──
  2929. if 'NMD_Decay_Class' in base_df.columns:
  2930. with st.expander("ORF Rules", expanded=False):
  2931. _render_facet_checkboxes(
  2932. _narrow(skip='decay')['NMD_Decay_Class'], key_prefix='f_decay',
  2933. order=_dom_decay,
  2934. label_fn=lambda v: DECAY_CLASS_EMOJI.get(v, v),
  2935. help=DECAY_CLASS_HELP,
  2936. )
  2937. # ── Kozak Context ──
  2938. if 'Kozak_Strength' in base_df.columns and base_df['Kozak_Strength'].notna().any():
  2939. with st.expander("Kozak Context", expanded=False):
  2940. _render_facet_checkboxes(
  2941. _narrow(skip='kozak')['Kozak_Strength'], key_prefix='f_kozak',
  2942. order=_dom_kozak,
  2943. help="Sequence-context strength around the start codon (Salk-discovered smORFs only).",
  2944. label_fn=lambda v: v.capitalize(),
  2945. )
  2946. # ── N-terminus Options ──
  2947. with st.expander("N-terminus Options", expanded=False):
  2948. _nv = _narrow(skip='nterm')
  2949. def _nterm_count(mode):
  2950. _m = _nterm_mask(_nv, mode)
  2951. return 0 if _m is None else int(_m.sum())
  2952. _nterm_help = (
  2953. "A tryptic peptide starting at position 1 or 2 is exactly the peptide "
  2954. "N-terminal Met excision of the parent ORF would produce, so on its own it "
  2955. "is not evidence the microprotein is real.\n\n"
  2956. "Non-N-terminal peptides only: keep microproteins with a tryptic peptide "
  2957. "starting at position 3 or later — evidence that survives removing the "
  2958. "M-excision peptides.\n\n"
  2959. "Nt-acetylated, substitution-distinct, or non-N-terminal: as above, but also "
  2960. "keep microproteins with an N-terminally acetylated peptide (co-translational, "
  2961. "marks a genuine N-terminus) OR a position-1/2 peptide whose own sequence carries "
  2962. "1-2 amino-acid substitutions vs. its matched UniProt isoform — a peptide that "
  2963. "differs from the canonical protein's sequence at that position can't be a bare "
  2964. "Met-excision fragment of it, so that's real evidence too.\n\n"
  2965. "Ticking both is the same as ticking the second — the first is a subset. "
  2966. "Both require tryptic evidence, so microproteins with no peptides are excluded."
  2967. )
  2968. for _i, _mode in enumerate(NTERM_MODES):
  2969. _count_checkbox(f"f_nterm_{_mode}", _mode, _nterm_count(_mode),
  2970. help=(_nterm_help if _i == 0 else None))
  2971. # ── Group by ──
  2972. _GROUP_BY_COLUMN_MAP = {
  2973. "smORF Type": "smORF_Group",
  2974. "Annotation Method": "Annotation_Status",
  2975. "Database Status": "_status_label",
  2976. "Start Codon": "Start_Codon",
  2977. }
  2978. with st.expander("Group by", expanded=False):
  2979. group_by_choice = st.selectbox(
  2980. "Group by", options=["None"] + list(_GROUP_BY_COLUMN_MAP.keys()),
  2981. key="f_group_by",
  2982. help="Reorder the table so rows sharing this value are grouped together; "
  2983. "shows a per-group count breakdown above the table.",
  2984. )
  2985. group_by_col = None
  2986. if group_by_choice != "None":
  2987. group_by_col = _GROUP_BY_COLUMN_MAP[group_by_choice]
  2988. if group_by_col == "_status_label":
  2989. filtered_df = filtered_df.copy()
  2990. filtered_df['_status_label'] = np.where(
  2991. _is_swiss_mask(filtered_df), 'Reviewed (Swiss-Prot)', 'Unreviewed (Salk/TrEMBL)'
  2992. )
  2993. if group_by_col not in filtered_df.columns:
  2994. group_by_col = None
  2995. st.markdown('<div class="sidebar-section-header">Legend</div>', unsafe_allow_html=True)
  2996. st.markdown("""
  2997. <div class="legend-container">
  2998. <div style="font-weight:600; margin-bottom:0.5rem; font-size:0.85rem;">Legend</div>
  2999. <div style="margin-bottom:0.35rem;"><span class="badge-swiss">Reviewed</span> Swiss-Prot</div>
  3000. <div><span class="badge-unreviewed">Unreviewed</span> Salk / TrEMBL</div>
  3001. </div>
  3002. """, unsafe_allow_html=True)
  3003. st.caption("Click links to view in UCSC Browser")
  3004. st.markdown(
  3005. "<div style='margin-top:0.6rem; font-size:0.8rem; color:#8da8b8;'>"
  3006. "<a href='https://github.com/brendan-miller-salk/brain-microprotein-atlas' "
  3007. "target='_blank' style='color:#74c2e1; text-decoration:none;'>"
  3008. "GitHub Repository</a></div>",
  3009. unsafe_allow_html=True,
  3010. )
  3011. # Every filter widget has now rendered — snapshot them so an entry-page
  3012. # detour on a later rerun can be undone.
  3013. _snapshot_filter_state()
  3014. # ── Command-center header with live metrics ──
  3015. total = len(filtered_df)
  3016. genes = filtered_df['Parent_Gene'].nunique() if 'Parent_Gene' in filtered_df.columns else 0
  3017. tmt_sig = int(filtered_df['TMT_Significant'].sum()) if 'TMT_Significant' in filtered_df.columns else 0
  3018. rna_sig = int(filtered_df['ROSMAP_Significant'].sum()) if 'ROSMAP_Significant' in filtered_df.columns else 0
  3019. swiss_count = int(_is_swiss_mask(filtered_df).sum())
  3020. noncan_count = total - swiss_count
  3021. header_slot.markdown(f"""
  3022. <div class="cmd-header">
  3023. <div class="cmd-top">
  3024. <div class="cmd-title">Brain Microprotein Dashboard</div>
  3025. <div class="cmd-subtitle">
  3026. Saghatelian Lab &middot; Salk Institute<br>
  3027. TMT-MS &middot; Ribo-seq &middot; Short/Long-Read RNA-seq
  3028. </div>
  3029. </div>
  3030. <div class="cmd-metrics">
  3031. <div class="cmd-stat st-total">
  3032. <div class="stat-label">Total</div>
  3033. <div class="stat-val">{total:,}</div>
  3034. <div class="stat-sub">{genes:,} genes</div>
  3035. </div>
  3036. <div class="cmd-stat st-swiss">
  3037. <div class="stat-label">Reviewed</div>
  3038. <div class="stat-val">{swiss_count:,}</div>
  3039. <div class="stat-sub">{swiss_count/total*100 if total else 0:.1f}% Swiss-Prot</div>
  3040. </div>
  3041. <div class="cmd-stat st-noncan">
  3042. <div class="stat-label">Unreviewed</div>
  3043. <div class="stat-val">{noncan_count:,}</div>
  3044. <div class="stat-sub">{noncan_count/total*100 if total else 0:.1f}% novel</div>
  3045. </div>
  3046. <div class="cmd-stat">
  3047. <div class="stat-label">smORF Groups</div>
  3048. <div class="stat-val">{filtered_df['smORF_Group'].nunique() if 'smORF_Group' in filtered_df.columns else 0}</div>
  3049. <div class="stat-sub">{filtered_df['smORF_Class'].nunique() if 'smORF_Class' in filtered_df.columns else 0} types</div>
  3050. </div>
  3051. <div class="cmd-stat">
  3052. <div class="stat-label">TMT Significant</div>
  3053. <div class="stat-val">{tmt_sig:,}</div>
  3054. <div class="stat-sub">q &lt; 0.2</div>
  3055. </div>
  3056. <div class="cmd-stat">
  3057. <div class="stat-label">RNA Significant</div>
  3058. <div class="stat-val">{rna_sig:,}</div>
  3059. <div class="stat-sub">padj &lt; 0.2</div>
  3060. </div>
  3061. </div>
  3062. </div>
  3063. """, unsafe_allow_html=True)
  3064. # ── Active-filter summary: show exactly which sidebar controls / presets are
  3065. # narrowing the current view, so users always know what is going on. ──
  3066. def _range_narrowed(col, rng):
  3067. if not rng or col not in unified_df.columns:
  3068. return False
  3069. _d = unified_df[col].dropna()
  3070. return (not _d.empty) and (rng[0] > _d.min() or rng[1] < _d.max())
  3071. _preset_labels = {
  3072. "lncrna": "lncRNA Microproteins",
  3073. "strong_graded": "Strong Graded Microproteins",
  3074. }
  3075. _active_chips = []
  3076. if hero_search and hero_search.strip():
  3077. _active_chips.append(("Search", hero_search.strip()))
  3078. if _qp in _preset_labels:
  3079. _active_chips.append(("Quick view", _preset_labels[_qp]))
  3080. if selected_groups:
  3081. _active_chips.append(("smORF type", ", ".join(map(str, selected_groups))))
  3082. if selected_downstream_sub:
  3083. _active_chips.append(("Downstream sub-type", ", ".join(map(str, selected_downstream_sub))))
  3084. if selected_decay:
  3085. _active_chips.append(("ORF Rules", ", ".join(map(str, selected_decay))))
  3086. if len(selected_status) == 1:
  3087. _active_chips.append(("Status", selected_status[0]))
  3088. if selected_quality:
  3089. _active_chips.append(("Spectra quality", ", ".join(map(str, selected_quality))))
  3090. if selected_min_peptides:
  3091. _active_chips.append(("Peptides", f"≥{min(selected_min_peptides)}"))
  3092. if selected_ribo:
  3093. _active_chips.append(("Ribosome coverage", ", ".join(map(str, selected_ribo))))
  3094. if selected_tmt_sig:
  3095. _active_chips.append(("TMT-MS", ", ".join(map(str, selected_tmt_sig))))
  3096. if selected_rna_sig:
  3097. _active_chips.append(("RNA", ", ".join(map(str, selected_rna_sig))))
  3098. if selected_codons:
  3099. _active_chips.append(("Start codon", ", ".join(map(str, selected_codons))))
  3100. if selected_kozak:
  3101. _active_chips.append(("Kozak strength", ", ".join(map(str, selected_kozak))))
  3102. if selected_shortstop:
  3103. _active_chips.append(("ShortStop", ", ".join(map(str, selected_shortstop))))
  3104. if selected_nterm:
  3105. _active_chips.append(("N-terminus", ", ".join(map(str, selected_nterm))))
  3106. if _range_narrowed('Protein_Length', length_range):
  3107. _active_chips.append(("Protein length", f"{int(length_range[0])}\u2013{int(length_range[1])} aa"))
  3108. if _range_narrowed('Unique_Spectral_Counts', spectral_range):
  3109. _active_chips.append(("Unique spectra", f"{int(spectral_range[0])}\u2013{int(spectral_range[1])}"))
  3110. if _range_narrowed('PhyloCSF_Score', phylocsf_range):
  3111. _active_chips.append(("PhyloCSF", f"{phylocsf_range[0]:.1f}\u2013{phylocsf_range[1]:.1f}"))
  3112. if _range_narrowed('UniProt_Annotation_Score', uniprot_score_range):
  3113. _active_chips.append(("UniProt score", f"{int(uniprot_score_range[0])}\u2013{int(uniprot_score_range[1])}\u2605"))
  3114. if group_by_col:
  3115. _active_chips.append(("Grouped by", group_by_choice))
  3116. _render_active_filter_bar(_active_chips, len(filtered_df), len(unified_df))
  3117. if len(filtered_df) == 0:
  3118. st.warning("No microproteins match your current filters. Try broadening the criteria.")
  3119. return
  3120. # ── Default sort: SAMs → Strong/Moderate spectra → lncRNA/uORF/dORF → high spec counts ──
  3121. _sort_df = filtered_df.copy()
  3122. # 1. ShortStop SAMs first (SAM-Intracellular, SAM-Secreted)
  3123. if 'ShortStop_Label' in _sort_df.columns:
  3124. _sort_df['_s1_sam'] = (~_sort_df['ShortStop_Label'].fillna('').str.startswith('SAM')).astype(int)
  3125. else:
  3126. _sort_df['_s1_sam'] = 1
  3127. # 2. Strong or Moderate spectra quality first
  3128. if '_Spectra_Quality' in _sort_df.columns:
  3129. _sort_df['_s2_quality'] = _sort_df['_Spectra_Quality'].map(
  3130. lambda q: 0 if q in ('Strong', 'Moderate') else 1
  3131. )
  3132. else:
  3133. _sort_df['_s2_quality'] = 1
  3134. # 3. lncRNA, uORF (Upstream), dORF (Downstream) smORF types first
  3135. _lnc_u_d_types = set(
  3136. ['lncRNA']
  3137. + SMORF_PARENT_GROUPS.get('Upstream', [])
  3138. + SMORF_PARENT_GROUPS.get('Downstream', [])
  3139. )
  3140. if 'smORF_Class' in _sort_df.columns:
  3141. _sort_df['_s3_type'] = (~_sort_df['smORF_Class'].isin(_lnc_u_d_types)).astype(int)
  3142. else:
  3143. _sort_df['_s3_type'] = 1
  3144. # 4. Highest spectral counts first
  3145. if 'Unique_Spectral_Counts' in _sort_df.columns:
  3146. _sort_df['_s4_spec'] = -_sort_df['Unique_Spectral_Counts'].fillna(0)
  3147. else:
  3148. _sort_df['_s4_spec'] = 0
  3149. _sort_cols = ['_s1_sam', '_s2_quality', '_s3_type', '_s4_spec']
  3150. if group_by_col and group_by_col in _sort_df.columns:
  3151. # Group-by is the primary sort key; existing tiers act as tiebreak
  3152. # within each group, so grouped rows still favor stronger evidence first.
  3153. _sort_cols = [group_by_col] + _sort_cols
  3154. filtered_df = _sort_df.sort_values(_sort_cols).drop(columns=[c for c in _sort_cols if c != group_by_col])
  3155. # ── Prepare display table ──
  3156. display_cols = ['sequence']
  3157. # NonATG Alternative Sequence(s): raw nonATG_predicted_sequence (a stringified
  3158. # Python list, e.g. "['MLWGR...', 'LLWGR...']") shown as-is, brackets included —
  3159. # blank for ATG-start / no-alternative-site microproteins.
  3160. if 'NonATG_Predicted_Sequence' in filtered_df.columns:
  3161. filtered_df['NonATG Alternative Sequence(s)'] = filtered_df['NonATG_Predicted_Sequence']
  3162. display_cols.append('NonATG Alternative Sequence(s)')
  3163. # Non-AUG alternative-initiation candidates, joined "codon@position" summary
  3164. # — Alt-Initiation tab only (see _render_results_table's column ordering).
  3165. if '_NonATG_Candidates_Display' in filtered_df.columns:
  3166. filtered_df['NonATG Candidates'] = filtered_df['_NonATG_Candidates_Display']
  3167. display_cols.append('NonATG Candidates')
  3168. if '_UCSC_HTML' in filtered_df.columns:
  3169. filtered_df = filtered_df.copy()
  3170. filtered_df['UCSC'] = filtered_df['_UCSC_HTML']
  3171. display_cols.append('UCSC')
  3172. elif 'UCSC_Link' in filtered_df.columns:
  3173. filtered_df = filtered_df.copy()
  3174. filtered_df['UCSC'] = filtered_df.apply(
  3175. lambda row: create_ucsc_link({'CLICK_UCSC': row['UCSC_Link']}, CUSTOM_UCSC_SESSION)
  3176. if pd.notna(row.get('UCSC_Link')) else None, axis=1
  3177. )
  3178. display_cols.append('UCSC')
  3179. column_mapping = {
  3180. 'Parent_Gene': 'Parent Gene',
  3181. 'smORF_Group': 'General smORF Type',
  3182. 'smORF_Display': 'smORF Subtype',
  3183. 'Protein_Length': 'Protein Length',
  3184. 'PhyloCSF_Score': 'PhyloCSF Score',
  3185. 'ShortStop_Label': 'ShortStop',
  3186. 'ShortStop_Score': 'ShortStop Score',
  3187. 'Annotation_Status': 'Annotation Method',
  3188. 'Unique_Spectral_Counts': 'Unique Spectral Counts (DDA)',
  3189. 'Razor_Spectral_Counts': 'Razor Counts (DDA)',
  3190. 'Start_Codon': 'Start Codon',
  3191. 'TMT_log2fc_50pct': 'TMT log2FC (50%)',
  3192. 'TMT_t_statistic_50pct': 'TMT t-stat (50%)',
  3193. 'TMT_df_50pct': 'TMT df (50%)',
  3194. 'TMT_conf_low_50pct': 'TMT CI low (50%)',
  3195. 'TMT_conf_high_50pct': 'TMT CI high (50%)',
  3196. 'TMT_cohens_d_50pct': "TMT Cohen's d (50%)",
  3197. 'TMT_pvalue_50pct': 'TMT p-val (50%)',
  3198. 'TMT_qvalue_50pct': 'TMT q-val (50%)',
  3199. 'TMT_log2fc_0pct': 'TMT log2FC (0%)',
  3200. 'TMT_t_statistic_0pct': 'TMT t-stat (0%)',
  3201. 'TMT_df_0pct': 'TMT df (0%)',
  3202. 'TMT_conf_low_0pct': 'TMT CI low (0%)',
  3203. 'TMT_conf_high_0pct': 'TMT CI high (0%)',
  3204. 'TMT_cohens_d_0pct': "TMT Cohen's d (0%)",
  3205. 'TMT_pvalue_0pct': 'TMT p-val (0%)',
  3206. 'TMT_qvalue_0pct': 'TMT q-val (0%)',
  3207. 'TMT_rate_control': 'MS Detect Control',
  3208. 'TMT_rate_ad': 'MS Detect AD',
  3209. 'ROSMAP_log2FC': 'ROSMAP log2FC',
  3210. 'ROSMAP_padj': 'ROSMAP padj',
  3211. 'MSBB_log2FC': 'MSBB log2FC',
  3212. 'MSBB_padj': 'MSBB padj',
  3213. 'Corr_MainORF_NonAD': 'ROSMAP Corr NonAD',
  3214. 'Corr_MainORF_AD': 'ROSMAP Corr AD',
  3215. 'Corr_MainORF_NonAD_MSBB': 'MSBB Corr NonAD',
  3216. 'Corr_MainORF_AD_MSBB': 'MSBB Corr AD',
  3217. 'RNA_LRT_Add_P': 'RNA_LRT_Add_P',
  3218. 'RNA_LRT_Int_P': 'RNA_LRT_Int_P',
  3219. 'RP3_Default': 'RP3 Default',
  3220. 'RP3_MM_Amb': 'RP3 MM+Amb',
  3221. 'RP3_Amb': 'RP3 Amb',
  3222. 'RP3_MM': 'RP3 MM',
  3223. 'RiboCode': 'RiboCode',
  3224. 'Psites_pct_frame0': 'P-site % Frame 0',
  3225. 'Psites_pct_frame1': 'P-site % Frame 1',
  3226. 'Psites_pct_frame2': 'P-site % Frame 2',
  3227. 'Psites_frame0_RPKM': 'P-site Frame 0 RPKM',
  3228. 'Tryptic_Peptides': 'Tryptic Peptides',
  3229. 'Tryptic_Protein_ID': 'Tryptic Protein ID',
  3230. 'Tryptic_Start_Positions': 'Tryptic Start Positions',
  3231. 'Tryptic_End_Positions': 'Tryptic End Positions',
  3232. 'Nt_Acetylated': 'Nt-Acetylated',
  3233. 'Nt_Acetyl_Peptides': 'Nt-Acetyl Peptides',
  3234. 'Nt_Acetyl_Total_PSMs': 'Nt-Acetyl PSMs',
  3235. 'Nt_Acetyl_Max_Fraction': 'Nt-Acetyl PSM Fraction',
  3236. 'BLAST_UniProt_Match': 'BLAST UniProt Match',
  3237. 'BLAST_Pct_Match': 'BLAST % Match',
  3238. 'BLAST_Aln_Length': 'BLAST Aln Length',
  3239. 'BLAST_Evalue': 'BLAST E-value',
  3240. 'BLAST_Bit': 'BLAST Bit Score',
  3241. 'UniProt_Annotation_Score': 'UniProt Annotation Score',
  3242. 'Kozak_Strength': 'Kozak Strength',
  3243. 'Kozak_Downstream_Strength': 'Kozak Downstream Strength',
  3244. 'Kozak_Class_Canonical': 'Kozak Class',
  3245. 'Kozak_Window': 'Kozak Window',
  3246. 'NonATG_Annotated_Codon': 'NonATG Codon',
  3247. 'NonATG_Is_Initiator': 'NonATG Valid Initiator',
  3248. 'NonATG_Has_Supported_Site': 'NonATG Has Supported Site',
  3249. 'NonATG_Context_Strength': 'NonATG Context Strength',
  3250. 'NonATG_N_Sites': 'NonATG Alt Sites Found',
  3251. 'NonATG_Best_Tier': 'Optimal Codon Tier',
  3252. }
  3253. for orig, display_name in column_mapping.items():
  3254. if orig in filtered_df.columns:
  3255. filtered_df[display_name] = filtered_df[orig]
  3256. display_cols.append(display_name)
  3257. # UniProt annotation score (gold stars; reviewed/TrEMBL only, blank otherwise)
  3258. if '_UniProt_Annotation_Display' in filtered_df.columns:
  3259. filtered_df['UniProt Annotation'] = filtered_df['_UniProt_Annotation_Display']
  3260. display_cols.append('UniProt Annotation')
  3261. # Mirror plot quality (use pre-computed column)
  3262. if '_Spectra_Quality' in filtered_df.columns:
  3263. filtered_df['Spectra Quality'] = filtered_df['_Spectra_Quality'].map(
  3264. lambda q: QUALITY_EMOJI.get(q, '\u2014')
  3265. )
  3266. display_cols.append('Spectra Quality')
  3267. # TMT significance tier (Tier 1 = strongest / Dark red … Tier 4 = Green)
  3268. if '_TMT_Tier' in filtered_df.columns:
  3269. filtered_df['TMT Tier'] = filtered_df['_TMT_Tier'].map(
  3270. lambda t: TMT_TIER_EMOJI.get(t, '\u2014')
  3271. )
  3272. display_cols.append('TMT Tier')
  3273. # RNA significance (Yellow = Exploratory, Green = Significant)
  3274. if '_RNA_Sig' in filtered_df.columns:
  3275. filtered_df['RNA Significance'] = filtered_df['_RNA_Sig'].map(
  3276. lambda t: RNA_SIG_EMOJI.get(t, '\u2014')
  3277. )
  3278. display_cols.append('RNA Significance')
  3279. # smORF decay class (NMD × main-ORF disruption); dots match the BED track
  3280. # colors. Raw NMD_Decay_Class stays untouched for the sidebar facet.
  3281. if 'NMD_Decay_Class' in filtered_df.columns:
  3282. filtered_df['ORF Rules'] = filtered_df['NMD_Decay_Class'].map(
  3283. lambda v: DECAY_CLASS_EMOJI.get(v, DECAY_CLASS_EMOJI['NA'])
  3284. )
  3285. display_cols.append('ORF Rules')
  3286. display_df = filtered_df[display_cols].copy()
  3287. # ── Format tryptic peptides as readable string (precomputed if available) ──
  3288. if 'Tryptic Peptides' in display_df.columns:
  3289. if '_Tryptic_Display' in filtered_df.columns:
  3290. display_df['Tryptic Peptides'] = filtered_df['_Tryptic_Display'].values
  3291. else:
  3292. def _fmt_peptides(val):
  3293. if not _not_na(val):
  3294. return ''
  3295. try:
  3296. peps = ast.literal_eval(str(val))
  3297. if isinstance(peps, str):
  3298. peps = [peps]
  3299. return ' \u00b7 '.join(str(p).strip() for p in peps)
  3300. except (ValueError, SyntaxError):
  3301. return str(val)
  3302. display_df['Tryptic Peptides'] = display_df['Tryptic Peptides'].map(_fmt_peptides)
  3303. # ── Mark Swiss-Prot rows: type columns get "Swiss-Prot" label, other
  3304. # annotation-only fields are non-applicable (em-dash). RP3 metrics
  3305. # DO apply to Swiss-Prot microproteins and should display normally. ──
  3306. _swiss_mask = None
  3307. if 'Database_All' in filtered_df.columns or 'Database' in filtered_df.columns:
  3308. _swiss_mask = _is_swiss_mask(filtered_df).values
  3309. # These columns are *not applicable* to Swiss-Prot reviewed entries.
  3310. _na_cols = ['ShortStop', 'Annotation Method']
  3311. for _c in _na_cols:
  3312. if _c in display_df.columns:
  3313. display_df[_c] = display_df[_c].fillna('').astype(str)
  3314. display_df.loc[_swiss_mask, _c] = '\u2014'
  3315. # smORF type columns get the "Swiss-Prot" badge for reviewed entries.
  3316. for _c in ('General smORF Type', 'smORF Subtype'):
  3317. if _c in display_df.columns:
  3318. display_df[_c] = display_df[_c].fillna('').astype(str)
  3319. display_df.loc[_swiss_mask, _c] = 'Swiss-Prot'
  3320. if 'ShortStop Score' in display_df.columns:
  3321. _ss = pd.to_numeric(display_df['ShortStop Score'], errors='coerce')
  3322. display_df['ShortStop Score'] = _ss.map(lambda v: f'{v:.3f}' if pd.notna(v) else '')
  3323. display_df.loc[_swiss_mask, 'ShortStop Score'] = '\u2014'
  3324. # ── Per-row background tint: muted teal (Swiss-Prot) vs muted orange (Unreviewed) ──
  3325. # NOTE: pandas Styler is too slow on ~6.5k rows; we use an emoji column
  3326. # ("DB") below for at-a-glance differentiation instead.
  3327. if _swiss_mask is not None:
  3328. display_df.insert(
  3329. 0,
  3330. 'DB',
  3331. np.where(_swiss_mask, '🔵 Swiss-Prot', '🟠 Unreviewed'),
  3332. )
  3333. # ── Group-by count breakdown (shown above the table when grouping is active) ──
  3334. if group_by_col and group_by_col in filtered_df.columns:
  3335. _group_counts = filtered_df[group_by_col].value_counts()
  3336. st.caption(
  3337. f"Grouped by {group_by_choice}: "
  3338. + " · ".join(f"{k}: {v:,}" for k, v in _group_counts.items())
  3339. )
  3340. # ── Data table (row click routes to the entry page) ──
  3341. _render_results_table(filtered_df, display_df)
  3342. # ── Downloads section ──
  3343. _render_downloads_section(filtered_df, display_df)
  3344. @st.fragment
  3345. def _render_results_table(filtered_df, display_df):
  3346. """The results table. Clicking a row navigates to that microprotein's entry
  3347. page (see _render_entry_page). Wrapped in @st.fragment so switching the
  3348. column view or opening the legend reruns only this block."""
  3349. # ── Column-group presets: limit horizontal scrolling by showing a curated
  3350. # subset of columns. Identity columns are always shown (and pinned left).
  3351. # Sequence columns are pushed to the far right everywhere EXCEPT
  3352. # Alt-Initiation, where they're the point and stay up front. NonATG
  3353. # Candidates is Alt-Initiation-only — hidden from every other view. ──
  3354. _sequence_cluster = ['sequence', 'NonATG Alternative Sequence(s)']
  3355. _alt_init_only = ['NonATG Candidates']
  3356. _core_identity = ['DB', 'UCSC', 'Parent Gene', 'General smORF Type', 'smORF Subtype',
  3357. 'ORF Rules']
  3358. # Alt-Initiation-only ordering: "NonATG Has Supported Site" sits as the 3rd
  3359. # column (right after "sequence"), matching its position in the detail
  3360. # card; "Optimal Codon Tier" follows, ahead of the alternative-sequence/
  3361. # candidates columns.
  3362. _identity_cols = (
  3363. ['DB', 'sequence', 'NonATG Has Supported Site', 'Optimal Codon Tier',
  3364. 'NonATG Alternative Sequence(s)']
  3365. + _alt_init_only
  3366. + [c for c in _core_identity if c != 'DB']
  3367. )
  3368. # Overview column order, authored end-to-end (identity columns included, and
  3369. # the sequence columns sitting mid-table rather than pushed right). Anything
  3370. # here that isn't in display_df is dropped, so this is safe to over-specify.
  3371. _OVERVIEW_ORDER = [
  3372. 'DB', 'UCSC', 'Parent Gene', 'General smORF Type', 'smORF Subtype',
  3373. 'Protein Length', 'Spectra Quality', 'PhyloCSF Score',
  3374. 'ShortStop', 'ShortStop Score', 'sequence', 'TMT Tier',
  3375. 'RNA Significance', 'Start Codon', 'ORF Rules', 'Kozak Strength',
  3376. 'NonATG Alternative Sequence(s)', 'UniProt Annotation',
  3377. ]
  3378. _column_groups = {
  3379. # Overview is the one view with a hand-authored full order (see
  3380. # _OVERVIEW_ORDER below): it interleaves the identity and sequence
  3381. # columns rather than taking the generic identity-first /
  3382. # sequence-last treatment, so it bypasses that assembly entirely.
  3383. 'Overview': [
  3384. 'General smORF Type', 'smORF Subtype', 'Protein Length',
  3385. 'Spectra Quality', 'PhyloCSF Score', 'ShortStop', 'ShortStop Score',
  3386. 'sequence', 'TMT Tier', 'RNA Significance', 'Start Codon',
  3387. 'ORF Rules', 'Kozak Strength', 'NonATG Alternative Sequence(s)',
  3388. 'UniProt Annotation',
  3389. ],
  3390. 'Proteomics': [
  3391. 'Unique Spectral Counts (DDA)', 'Razor Counts (DDA)',
  3392. 'TMT log2FC (50%)', 'TMT t-stat (50%)', 'TMT df (50%)',
  3393. 'TMT CI low (50%)', 'TMT CI high (50%)', "TMT Cohen's d (50%)",
  3394. 'TMT p-val (50%)', 'TMT q-val (50%)',
  3395. 'TMT log2FC (0%)', 'TMT t-stat (0%)', 'TMT df (0%)',
  3396. 'TMT CI low (0%)', 'TMT CI high (0%)', "TMT Cohen's d (0%)",
  3397. 'TMT p-val (0%)', 'TMT q-val (0%)',
  3398. 'MS Detect Control', 'MS Detect AD', 'TMT Tier', 'Spectra Quality',
  3399. 'Tryptic Peptides', 'Tryptic Protein ID',
  3400. 'Tryptic Start Positions', 'Tryptic End Positions',
  3401. 'Nt-Acetylated', 'Nt-Acetyl Peptides', 'Nt-Acetyl PSMs',
  3402. 'Nt-Acetyl PSM Fraction',
  3403. ],
  3404. 'Transcriptomics': [
  3405. 'ROSMAP log2FC', 'MSBB log2FC',
  3406. 'ROSMAP padj', 'MSBB padj',
  3407. 'ROSMAP Corr NonAD', 'MSBB Corr NonAD',
  3408. 'ROSMAP Corr AD', 'MSBB Corr AD',
  3409. 'RNA_LRT_Add_P', 'RNA_LRT_Int_P', 'RNA Significance',
  3410. ],
  3411. 'Ribo-seq': [
  3412. 'RiboCode', 'P-site Frame 0 RPKM',
  3413. 'P-site % Frame 0', 'P-site % Frame 1', 'P-site % Frame 2',
  3414. 'RP3 Default', 'RP3 MM+Amb', 'RP3 Amb', 'RP3 MM',
  3415. ],
  3416. 'Homology (BLAST)': [
  3417. 'BLAST UniProt Match', 'BLAST % Match', 'BLAST Aln Length',
  3418. 'BLAST E-value', 'BLAST Bit Score',
  3419. ],
  3420. 'Alt-Initiation': [
  3421. 'General smORF Type', 'smORF Subtype', 'ORF Rules', 'Start Codon',
  3422. 'Kozak Strength', 'Kozak Downstream Strength', 'Kozak Class', 'Kozak Window',
  3423. 'NonATG Codon', 'NonATG Valid Initiator', 'NonATG Has Supported Site',
  3424. 'NonATG Context Strength', 'NonATG Alt Sites Found',
  3425. ],
  3426. 'All columns': None,
  3427. }
  3428. _preset = st.segmented_control(
  3429. 'Column view',
  3430. options=list(_column_groups.keys()),
  3431. default='Overview',
  3432. help='Switch between curated column sets to avoid scrolling. '
  3433. 'Identity columns (DB, UCSC, Parent Gene, smORF type/subtype) are shown in every '
  3434. 'view; sequence columns (Sequence, NonATG Alternative Sequence(s)) are pushed to '
  3435. 'the far right, except in Alt-Initiation where they stay up front. NonATG '
  3436. 'Candidates only appears in the Alt-Initiation view.',
  3437. ) or 'Overview'
  3438. _group = _column_groups.get(_preset)
  3439. if _preset == 'Overview':
  3440. _column_order = [c for c in _OVERVIEW_ORDER if c in display_df.columns]
  3441. elif _preset == 'Alt-Initiation':
  3442. # Alt-Initiation is only meaningful for nonATG-start smORFs — ATG-start
  3443. # rows have no alternative-initiation candidates to show.
  3444. if 'Start Codon' in display_df.columns:
  3445. display_df = display_df[display_df['Start Codon'] == 'nonATG']
  3446. # Sequence info (and NonATG Candidates) stays up front — it's the whole
  3447. # point of this view.
  3448. _ordered = _identity_cols + [c for c in _group if c not in _identity_cols]
  3449. _column_order = [c for c in _ordered if c in display_df.columns]
  3450. elif _group is None:
  3451. # 'All columns' — core identity first, everything else in its natural
  3452. # order, sequence cluster pushed to the very end. NonATG Candidates is
  3453. # excluded entirely outside Alt-Initiation.
  3454. _rest = [c for c in display_df.columns
  3455. if c not in _core_identity and c not in _sequence_cluster
  3456. and c not in _alt_init_only]
  3457. _ordered = _core_identity + _rest + _sequence_cluster
  3458. _column_order = [c for c in _ordered if c in display_df.columns]
  3459. else:
  3460. _ordered = (_core_identity + [c for c in _group if c not in _core_identity]
  3461. + _sequence_cluster)
  3462. _column_order = [c for c in _ordered if c in display_df.columns]
  3463. # ── Column legend: describes whatever columns are visible in the *current*
  3464. # tab, so it updates each time a different Column view is selected. ──
  3465. with st.expander("Explain the column/variables to me", expanded=False):
  3466. _legend_cols = _column_order if _column_order is not None else list(display_df.columns)
  3467. st.markdown("\n".join(
  3468. f"- **{c}** — {COLUMN_DESCRIPTIONS[c]}"
  3469. for c in _legend_cols if c in COLUMN_DESCRIPTIONS
  3470. ) or "No column descriptions available for this view.")
  3471. # The row-select box in the leftmost column is the only in-place control the
  3472. # grid offers — data-cell clicks do nothing, and st.column_config.LinkColumn
  3473. # always opens a new tab (it calls window.open(url, "_blank") internally), so
  3474. # a link column can't be used for same-tab navigation. Hence the signpost.
  3475. st.markdown(
  3476. '<div class="open-hint">'
  3477. '<span class="open-hint-arrow">&#11013;</span>'
  3478. '<span>Tick the <strong>select box</strong> at the start of any row to open that '
  3479. 'microprotein&rsquo;s full entry page &mdash; annotation, evidence by platform, '
  3480. 'figures and sequence.</span>'
  3481. '</div>',
  3482. unsafe_allow_html=True,
  3483. )
  3484. selection = st.dataframe(
  3485. display_df,
  3486. width='stretch',
  3487. hide_index=True,
  3488. on_select="rerun",
  3489. selection_mode="single-row",
  3490. # Nonce-suffixed key: returning from an entry page bumps the nonce so the
  3491. # widget is recreated without its stale row selection (see
  3492. # _back_to_results). Without this the table would immediately navigate
  3493. # back to the entry we just left.
  3494. key=f"results_table_{st.session_state.get('_table_nonce', 0)}",
  3495. column_order=_column_order,
  3496. column_config={
  3497. 'DB': st.column_config.TextColumn(
  3498. 'DB', help='🔵 Swiss-Prot (reviewed) · 🟠 Unreviewed', width='small'
  3499. ),
  3500. 'UCSC': st.column_config.LinkColumn(
  3501. 'UCSC', help='View in UCSC Genome Browser', display_text='View'
  3502. ),
  3503. 'sequence': st.column_config.TextColumn('Sequence', max_chars=50),
  3504. 'NonATG Alternative Sequence(s)': st.column_config.TextColumn(
  3505. 'NonATG Alternative Sequence(s)', max_chars=50,
  3506. help='Predicted sequence(s) from alternative (non-AUG) initiation candidates, '
  3507. 'raw list format (e.g. [\'MLWGR...\', \'LLWGR...\']); blank for ATG-start '
  3508. 'microproteins or non-ATG starts with no compatible alternative site.'
  3509. ),
  3510. 'Parent Gene': st.column_config.TextColumn('Parent Gene', max_chars=15),
  3511. 'General smORF Type': st.column_config.TextColumn('General smORF Type', max_chars=20),
  3512. 'smORF Subtype': st.column_config.TextColumn('smORF Subtype', max_chars=15),
  3513. 'ORF Rules': st.column_config.TextColumn(
  3514. 'ORF Rules', max_chars=27, help=DECAY_CLASS_HELP),
  3515. 'ShortStop Score': st.column_config.TextColumn('ShortStop Score', help='ShortStop ML confidence score (0-1)', max_chars=10),
  3516. 'PhyloCSF Score': st.column_config.NumberColumn('PhyloCSF Score', help='PhyloCSF evolutionary conservation score', format='%.2f'),
  3517. 'Unique Spectral Counts (DDA)': st.column_config.NumberColumn('Unique Spectral Counts (DDA)', help='Number of unique mass spectrometry spectral counts', format='%d'),
  3518. 'Razor Counts (DDA)': st.column_config.NumberColumn('Razor Counts (DDA)', help='Total razor spectral counts (shared peptides assigned to this protein)', format='%d'),
  3519. 'Protein Length': st.column_config.NumberColumn('Protein Length', help='Length of protein in amino acids', format='%d'),
  3520. 'Annotation Method': st.column_config.TextColumn('Annotation Method', help='Method used for annotation (MS = Mass Spectrometry, etc.)', max_chars=10),
  3521. 'TMT log2FC (50%)': st.column_config.NumberColumn('TMT log2FC (50%)', help='TMT log2 fold-change (AD vs Control) — 50% missing threshold', format='%.3f'),
  3522. 'TMT t-stat (50%)': st.column_config.NumberColumn('TMT t-stat (50%)', help='TMT t-statistic (AD vs Control) — 50% missing threshold', format='%.3f'),
  3523. 'TMT df (50%)': st.column_config.NumberColumn('TMT df (50%)', help='TMT t-test degrees of freedom — 50% missing threshold', format='%.1f'),
  3524. 'TMT CI low (50%)': st.column_config.NumberColumn('TMT CI low (50%)', help='TMT 95% CI lower bound on log2 fold-change — 50% missing threshold', format='%.3f'),
  3525. 'TMT CI high (50%)': st.column_config.NumberColumn('TMT CI high (50%)', help='TMT 95% CI upper bound on log2 fold-change — 50% missing threshold', format='%.3f'),
  3526. "TMT Cohen's d (50%)": st.column_config.NumberColumn("TMT Cohen's d (50%)", help='TMT Cohen\'s d effect size (AD vs Control) — 50% missing threshold', format='%.3f'),
  3527. 'TMT p-val (50%)': st.column_config.NumberColumn('TMT p-val (50%)', help='TMT raw p-value — 50% missing threshold', format='%.4f'),
  3528. 'TMT q-val (50%)': st.column_config.NumberColumn('TMT q-val (50%)', help='TMT q-value (BH-adjusted) — 50% missing threshold', format='%.4f'),
  3529. 'TMT log2FC (0%)': st.column_config.NumberColumn('TMT log2FC (0%)', help='TMT log2 fold-change (AD vs Control) — 0% missing threshold (stringent)', format='%.3f'),
  3530. 'TMT t-stat (0%)': st.column_config.NumberColumn('TMT t-stat (0%)', help='TMT t-statistic (AD vs Control) — 0% missing threshold (stringent)', format='%.3f'),
  3531. 'TMT df (0%)': st.column_config.NumberColumn('TMT df (0%)', help='TMT t-test degrees of freedom — 0% missing threshold (stringent)', format='%.1f'),
  3532. 'TMT CI low (0%)': st.column_config.NumberColumn('TMT CI low (0%)', help='TMT 95% CI lower bound on log2 fold-change — 0% missing threshold (stringent)', format='%.3f'),
  3533. 'TMT CI high (0%)': st.column_config.NumberColumn('TMT CI high (0%)', help='TMT 95% CI upper bound on log2 fold-change — 0% missing threshold (stringent)', format='%.3f'),
  3534. "TMT Cohen's d (0%)": st.column_config.NumberColumn("TMT Cohen's d (0%)", help='TMT Cohen\'s d effect size (AD vs Control) — 0% missing threshold (stringent)', format='%.3f'),
  3535. 'TMT p-val (0%)': st.column_config.NumberColumn('TMT p-val (0%)', help='TMT raw p-value — 0% missing threshold (stringent)', format='%.4f'),
  3536. 'TMT q-val (0%)': st.column_config.NumberColumn('TMT q-val (0%)', help='TMT q-value (BH-adjusted) — 0% missing threshold (stringent)', format='%.4f'),
  3537. 'MS Detect Control': st.column_config.NumberColumn('MS Detect Control', help='MS detection rate in Control donors', format='%.3f'),
  3538. 'MS Detect AD': st.column_config.NumberColumn('MS Detect AD', help='MS detection rate in AD donors', format='%.3f'),
  3539. 'ROSMAP log2FC': st.column_config.NumberColumn('ROSMAP log2FC', help='ROSMAP short-read RNA-seq log2 fold-change (AD vs Control)', format='%.3f'),
  3540. 'ROSMAP padj': st.column_config.NumberColumn('ROSMAP padj', help='ROSMAP short-read RNA-seq adjusted p-value', format='%.4f'),
  3541. 'MSBB log2FC': st.column_config.NumberColumn('MSBB log2FC', help='MSBB short-read RNA-seq log2 fold-change (AD vs Control)', format='%.3f'),
  3542. 'MSBB padj': st.column_config.NumberColumn('MSBB padj', help='MSBB short-read RNA-seq adjusted p-value', format='%.4f'),
  3543. 'ROSMAP Corr NonAD': st.column_config.NumberColumn('ROSMAP Corr NonAD', help='ROSMAP: correlation between Main ORF and smORF transcripts (non-AD donors)', format='%.3f'),
  3544. 'ROSMAP Corr AD': st.column_config.NumberColumn('ROSMAP Corr AD', help='ROSMAP: correlation between Main ORF and smORF transcripts (AD donors)', format='%.3f'),
  3545. 'MSBB Corr NonAD': st.column_config.NumberColumn('MSBB Corr NonAD', help='MSBB: correlation between Main ORF and smORF transcripts (non-AD donors)', format='%.3f'),
  3546. 'MSBB Corr AD': st.column_config.NumberColumn('MSBB Corr AD', help='MSBB: correlation between Main ORF and smORF transcripts (AD donors)', format='%.3f'),
  3547. 'RNA_LRT_Add_P': st.column_config.NumberColumn('RNA_LRT_Add_P', help='ROSMAP RNA-seq LRT additive model p-value', format='%.4f'),
  3548. 'RNA_LRT_Int_P': st.column_config.NumberColumn('RNA_LRT_Int_P', help='ROSMAP RNA-seq LRT interaction model p-value', format='%.4f'),
  3549. 'RP3 Default': st.column_config.TextColumn('RP3 Default', help='RP3 default pipeline result', max_chars=10),
  3550. 'RP3 MM+Amb': st.column_config.TextColumn('RP3 MM+Amb', help='RP3 multi-mapping + ambiguous result', max_chars=10),
  3551. 'RP3 Amb': st.column_config.TextColumn('RP3 Amb', help='RP3 ambiguous reads result', max_chars=10),
  3552. 'RP3 MM': st.column_config.TextColumn('RP3 MM', help='RP3 multi-mapping result', max_chars=10),
  3553. 'RiboCode': st.column_config.TextColumn('RiboCode', help='RiboCode ORF detection result', max_chars=10),
  3554. 'P-site % Frame 0': st.column_config.NumberColumn('P-site % Frame 0', help='Percent of ORF P-sites in frame 0 (in-frame codon position)', format='%.1f'),
  3555. 'P-site % Frame 1': st.column_config.NumberColumn('P-site % Frame 1', help='Percent of ORF P-sites in frame 1', format='%.1f'),
  3556. 'P-site % Frame 2': st.column_config.NumberColumn('P-site % Frame 2', help='Percent of ORF P-sites in frame 2', format='%.1f'),
  3557. 'P-site Frame 0 RPKM': st.column_config.NumberColumn('P-site Frame 0 RPKM', help='Frame-0 P-sites per kb of CDS per million P-sites', format='%.3f'),
  3558. 'Tryptic Peptides': st.column_config.TextColumn('Tryptic Peptides', help='Tryptic peptide sequences', max_chars=60),
  3559. 'Tryptic Protein ID': st.column_config.TextColumn('Protein ID', help='Protein ID associated with the tryptic peptides', max_chars=15),
  3560. 'Tryptic Start Positions': st.column_config.TextColumn('Start Positions', help='Start positions of tryptic peptides', max_chars=30),
  3561. 'Tryptic End Positions': st.column_config.TextColumn('End Positions', help='End positions of tryptic peptides', max_chars=30),
  3562. 'Nt-Acetylated': st.column_config.CheckboxColumn('Nt-Acetylated', help='At least one tryptic peptide carries an N-terminal acetyl mark — direct evidence that this is a genuine protein N-terminus'),
  3563. 'Nt-Acetyl Peptides': st.column_config.TextColumn('Nt-Acetyl Peptides', help='Tryptic peptide sequence(s) observed with an N-terminal acetyl mark', max_chars=60),
  3564. 'Nt-Acetyl PSMs': st.column_config.NumberColumn('Nt-Acetyl PSMs', help='Total Nt-acetylated PSMs summed across the Nt-acetylated peptides', format='%d'),
  3565. 'Nt-Acetyl PSM Fraction': st.column_config.NumberColumn('Nt-Acetyl PSM Fraction', help='Highest per-peptide fraction of that peptide’s PSMs that carry the Nt-acetyl mark (1.0 = every PSM acetylated)', format='%.3f'),
  3566. 'Spectra Quality': st.column_config.TextColumn('Spectra Quality', help='Best Prosit spectral-match confidence tier (from the master Confidence column)', max_chars=20),
  3567. 'BLAST UniProt Match': st.column_config.TextColumn('BLAST UniProt Match', help='UniProt accession of the best BLASTp hit', max_chars=15),
  3568. 'BLAST % Match': st.column_config.NumberColumn('BLAST % Match', help='BLASTp percent identity to the best UniProt hit', format='%.1f'),
  3569. 'BLAST Aln Length': st.column_config.NumberColumn('BLAST Aln Length', help='BLASTp alignment length (residues)', format='%d'),
  3570. 'BLAST E-value': st.column_config.NumberColumn('BLAST E-value', help='BLASTp expectation value of the best UniProt hit', format='%.2e'),
  3571. 'BLAST Bit Score': st.column_config.NumberColumn('BLAST Bit Score', help='BLASTp bit score of the best UniProt hit', format='%.1f'),
  3572. 'UniProt Annotation': st.column_config.TextColumn('UniProt Annotation', help='UniProt curation annotation score (1\u20135\u2605) for reviewed / TrEMBL entries; blank for novel microproteins not in UniProt', max_chars=8),
  3573. 'UniProt Annotation Score': st.column_config.NumberColumn('UniProt Annotation Score', help='Numeric UniProt annotation score (1\u20135); use to sort by curation depth', format='%d'),
  3574. 'Kozak Strength': st.column_config.TextColumn('Kozak Strength', help='Full Kozak consensus context strength around the start codon (Salk smORFs only)', max_chars=10),
  3575. 'Kozak Downstream Strength': st.column_config.TextColumn('Kozak Downstream Strength', help='Kozak context strength considering only the +4 downstream position', max_chars=10),
  3576. 'Kozak Class': st.column_config.TextColumn('Kozak Class', help='Canonical Kozak strength classification', max_chars=10),
  3577. 'Kozak Window': st.column_config.TextColumn('Kozak Window', help='Local sequence context around the start codon (lowercase = flanking, uppercase = start + downstream)', max_chars=20),
  3578. 'NonATG Codon': st.column_config.TextColumn('NonATG Codon', help='The originally annotated (non-ATG) start codon', max_chars=6),
  3579. 'NonATG Valid Initiator': st.column_config.TextColumn('NonATG Valid Initiator', help='Whether the annotated codon is itself a plausible translation initiator', max_chars=6),
  3580. 'NonATG Has Supported Site': st.column_config.TextColumn('NonATG Has Supported Site', help='Whether current non-canonical initiation evidence/models support this annotated start actually being used', max_chars=6),
  3581. 'NonATG Context Strength': st.column_config.TextColumn('NonATG Context Strength', help='Kozak context strength around the annotated non-ATG codon', max_chars=10),
  3582. 'NonATG Alt Sites Found': st.column_config.NumberColumn('NonATG Alt Sites Found', help='Number of peptide-compatible alternative initiation sites found by scanning', format='%d'),
  3583. 'Optimal Codon Tier': st.column_config.TextColumn('Optimal Codon Tier', help='Strongest evidence tier among the alternative initiation candidates', max_chars=40),
  3584. 'NonATG Candidates': st.column_config.TextColumn('NonATG Candidates', help='Alternative initiation candidates as codon@position', max_chars=40),
  3585. }
  3586. )
  3587. # ── Row click: navigate to that microprotein's entry page ──
  3588. # display_df.index[...] maps the widget's *positional* selection back to the
  3589. # label shared with filtered_df — the two frames diverge whenever a column
  3590. # view sub-filters display_df (e.g. Alt-Initiation), so never index
  3591. # filtered_df positionally here.
  3592. selected_rows = selection.selection.rows if selection else []
  3593. if selected_rows:
  3594. _row = filtered_df.loc[display_df.index[selected_rows[0]]]
  3595. st.query_params[ENTRY_QUERY_PARAM] = _entry_id(_row.get('sequence', ''))
  3596. # scope="app": this function is a fragment, so the default fragment-scoped
  3597. # rerun would never reach the routing check in main().
  3598. st.rerun(scope="app")
  3599. else:
  3600. st.markdown('<div class="detail-panel"><div class="detail-prompt">Tick a row\'s select box above to open that microprotein\'s full entry page</div></div>', unsafe_allow_html=True)
  3601. @st.cache_data
  3602. def _read_static_file(path_str):
  3603. """Read a static file's bytes for download (cached to avoid re-reads)."""
  3604. p = Path(path_str)
  3605. if p.exists():
  3606. return p.read_bytes()
  3607. return None
  3608. def _fmt_size(path):
  3609. """Return human-readable file size string, or empty string if missing."""
  3610. p = Path(path)
  3611. if not p.exists():
  3612. return ""
  3613. n = p.stat().st_size
  3614. for unit in ("B", "KB", "MB", "GB"):
  3615. if n < 1024:
  3616. return f"{n:.0f} {unit}"
  3617. n /= 1024
  3618. return f"{n:.1f} GB"
  3619. def _filtered_fasta(df):
  3620. """Generate a FASTA string from the filtered dataframe's sequences."""
  3621. lines = []
  3622. for _, row in df.iterrows():
  3623. seq = row.get("sequence", "")
  3624. if not seq or pd.isna(seq):
  3625. continue
  3626. uid = row.get("UniProt_ID", "") or row.get("smORF_Coordinates", "") or "smORF"
  3627. gene = row.get("Parent_Gene", "")
  3628. header = f">{uid}"
  3629. if gene and pd.notna(gene):
  3630. header += f"|{gene}"
  3631. lines.append(header)
  3632. seq = str(seq).strip()
  3633. for i in range(0, len(seq), 60):
  3634. lines.append(seq[i:i + 60])
  3635. return "\n".join(lines)
  3636. def _render_downloads_section(filtered_df, display_df):
  3637. """Full downloads section rendered below the main table."""
  3638. with st.expander("Downloads", expanded=False):
  3639. # ── 1. Current View ──────────────────────────────────────────────────
  3640. st.markdown(
  3641. "<div style='font-size:0.8rem; font-weight:600; color:#8da8b8; "
  3642. "text-transform:uppercase; letter-spacing:0.05em; margin-bottom:0.4rem;'>"
  3643. "Current View</div>",
  3644. unsafe_allow_html=True,
  3645. )
  3646. cv1, cv2, _ = st.columns([2, 2, 3])
  3647. with cv1:
  3648. st.download_button(
  3649. label=f"Table (.csv) — {len(display_df):,} rows",
  3650. data=display_df.to_csv(index=False),
  3651. file_name=f"brain_microproteins_filtered_{len(display_df)}.csv",
  3652. mime="text/csv",
  3653. use_container_width=True,
  3654. key="dl_filtered_csv",
  3655. )
  3656. with cv2:
  3657. fasta_data = _filtered_fasta(filtered_df)
  3658. fasta_count = fasta_data.count(">")
  3659. st.download_button(
  3660. label=f"Sequences (.fasta) — {fasta_count:,} entries",
  3661. data=fasta_data,
  3662. file_name=f"brain_microproteins_filtered_{fasta_count}.fasta",
  3663. mime="text/plain",
  3664. use_container_width=True,
  3665. disabled=fasta_count == 0,
  3666. key="dl_filtered_fasta",
  3667. )
  3668. st.divider()
  3669. # ── 2. Genome Annotation ─────────────────────────────────────────────
  3670. st.markdown(
  3671. "<div style='font-size:0.8rem; font-weight:600; color:#8da8b8; "
  3672. "text-transform:uppercase; letter-spacing:0.05em; margin-bottom:0.4rem;'>"
  3673. "Genome Annotation — full dataset</div>",
  3674. unsafe_allow_html=True,
  3675. )
  3676. _GENOME_FILES = [
  3677. (
  3678. "Unreviewed_Brain_Microproteins.fasta",
  3679. "FASTA",
  3680. "Amino acid sequences for all unreviewed brain microproteins",
  3681. "text/plain",
  3682. ),
  3683. (
  3684. "Unreviewed_Brain_Microproteins_CDS_Absent_from_UniProt.bed",
  3685. "BED",
  3686. "CDS genomic coordinates (hg38) — absent from UniProt",
  3687. "text/plain",
  3688. ),
  3689. (
  3690. "Unreviewed_Brain_Microproteins_Absent_from_UniProt.gtf",
  3691. "GTF",
  3692. "Gene transfer format annotations — absent from UniProt",
  3693. "text/plain",
  3694. ),
  3695. (
  3696. "Unreviewed_Brain_Microproteins_IDs.txt",
  3697. "TXT",
  3698. "Protein accession ID list",
  3699. "text/plain",
  3700. ),
  3701. (
  3702. "Unreviewed_Brain_Microproteins_genomic_coordinates.txt",
  3703. "TXT",
  3704. "Genomic coordinate list (chr:start-end)",
  3705. "text/plain",
  3706. ),
  3707. (
  3708. "Unreviewed_Brain_Microproteins_mapping_coordinates_to_sequences.tsv",
  3709. "TSV",
  3710. "Sequence ↔ coordinate mapping table",
  3711. "text/tab-separated-values",
  3712. ),
  3713. ]
  3714. for filename, fmt_tag, description, mime in _GENOME_FILES:
  3715. fpath = GENOME_FILES_DIR / filename
  3716. size_str = _fmt_size(fpath)
  3717. data = _read_static_file(str(fpath))
  3718. col_tag, col_name, col_size, col_btn = st.columns([1, 5, 1, 1])
  3719. with col_tag:
  3720. st.markdown(
  3721. f"<span style='background:#1e3a4f; color:#74c2e1; font-size:0.7rem; "
  3722. f"font-weight:700; padding:2px 6px; border-radius:4px; "
  3723. f"letter-spacing:0.04em;'>{fmt_tag}</span>",
  3724. unsafe_allow_html=True,
  3725. )
  3726. with col_name:
  3727. st.markdown(
  3728. f"<span style='font-size:0.85rem;'>{filename}</span> \n"
  3729. f"<span style='font-size:0.75rem; color:#8da8b8;'>{description}</span>",
  3730. unsafe_allow_html=True,
  3731. )
  3732. with col_size:
  3733. st.markdown(
  3734. f"<span style='font-size:0.75rem; color:#8da8b8;'>{size_str}</span>",
  3735. unsafe_allow_html=True,
  3736. )
  3737. with col_btn:
  3738. st.download_button(
  3739. label="Download",
  3740. data=data if data else b"",
  3741. file_name=filename,
  3742. mime=mime,
  3743. use_container_width=True,
  3744. disabled=data is None,
  3745. key=f"dl_genome_{filename}",
  3746. )
  3747. # Large combined GTF — repository only
  3748. large_gtf = GENOME_FILES_DIR / "Ensembl_and_Unreviewed_Brain_Microproteins.gtf"
  3749. large_size = _fmt_size(large_gtf)
  3750. col_tag, col_name, col_size, col_btn = st.columns([1, 5, 1, 1])
  3751. with col_tag:
  3752. st.markdown(
  3753. "<span style='background:#1e3a4f; color:#74c2e1; font-size:0.7rem; "
  3754. "font-weight:700; padding:2px 6px; border-radius:4px; "
  3755. "letter-spacing:0.04em;'>GTF</span>",
  3756. unsafe_allow_html=True,
  3757. )
  3758. with col_name:
  3759. st.markdown(
  3760. "<span style='font-size:0.85rem;'>Ensembl_and_Unreviewed_Brain_Microproteins.gtf</span> \n"
  3761. "<span style='font-size:0.75rem; color:#8da8b8;'>Combined Ensembl + unreviewed annotations "
  3762. "— available via project repository</span>",
  3763. unsafe_allow_html=True,
  3764. )
  3765. with col_size:
  3766. st.markdown(
  3767. f"<span style='font-size:0.75rem; color:#8da8b8;'>{large_size}</span>",
  3768. unsafe_allow_html=True,
  3769. )
  3770. with col_btn:
  3771. st.markdown(
  3772. "<span style='font-size:0.75rem; color:#8da8b8;'>too large</span>",
  3773. unsafe_allow_html=True,
  3774. )
  3775. def _fmt(val, fmt_str=None):
  3776. """Format a value for display, returning '—' for NaN/None."""
  3777. if val is None or (isinstance(val, float) and pd.isna(val)):
  3778. return '—'
  3779. if fmt_str:
  3780. try:
  3781. return fmt_str % val
  3782. except (TypeError, ValueError):
  3783. return str(val)
  3784. return str(val)
  3785. def _not_na(val):
  3786. """Return True if val is not None/NaN."""
  3787. if val is None:
  3788. return False
  3789. if isinstance(val, float) and pd.isna(val):
  3790. return False
  3791. return True
  3792. def _sig_class_fc(val, sig_val=None, threshold=0.2):
  3793. """CSS class for fold-change direction coloring (muted when not significant)."""
  3794. if not _not_na(val):
  3795. return 'sig-na'
  3796. if sig_val is not None and _not_na(sig_val) and sig_val >= threshold:
  3797. return 'sig-ns'
  3798. return 'sig-up' if val > 0 else 'sig-down' if val < 0 else ''
  3799. def _sig_class_pval(val, threshold):
  3800. """CSS class for p-value significance coloring."""
  3801. if not _not_na(val):
  3802. return 'sig-na'
  3803. return 'sig-yes' if val < threshold else 'sig-ns'
  3804. # =============================================================================
  3805. # scRNA CELL-TYPE ENRICHMENT PANEL (single-microprotein heatmap + sig bars)
  3806. # =============================================================================
  3807. @st.cache_data(show_spinner=False)
  3808. def load_scrna_celltype_table(_all_path=str(SCRNA_ALL_CELLTYPES_FILE),
  3809. _sig_path=str(SCRNA_SUMMARY_FILE)):
  3810. """Per-(microprotein, cell type) enrichment rows, trimmed to the plotted
  3811. columns and indexed by sequence.
  3812. Returns `(df, complete)`. `complete` is True when the all-cell-types export
  3813. is present, i.e. the frame holds every cell type that was tested and the
  3814. `significant` column separates the hits. When only the significant-only
  3815. summary is shipped, everything in the frame is significant and `complete` is
  3816. False, so the panel can say that cell types are missing rather than draw a
  3817. partial picture as if it were the whole one.
  3818. `sequence` is the join key the whole dashboard uses and it matches this
  3819. table exactly (every sequence in it is in the master).
  3820. """
  3821. cols = ['sequence', 'gene', 'celltype', 'cell_type_general',
  3822. 'logFC', 'logCPM', 'p_adj.glb', 'p_adj.loc', 'significant']
  3823. df, complete = None, False
  3824. p_all = Path(_all_path)
  3825. if p_all.exists():
  3826. try:
  3827. df = pd.read_csv(p_all, usecols=lambda c: c in cols)
  3828. complete = True
  3829. except Exception:
  3830. df = None
  3831. if df is None:
  3832. p_sig = Path(_sig_path)
  3833. if not p_sig.exists():
  3834. return None, False
  3835. try:
  3836. df = pd.read_csv(p_sig, low_memory=False, usecols=lambda c: c in cols)
  3837. except Exception:
  3838. return None, False
  3839. # This file is the p_adj.glb < 0.05 subset by construction.
  3840. df['significant'] = True
  3841. if 'sequence' not in df.columns:
  3842. return None, False
  3843. for c in ('logFC', 'logCPM', 'p_adj

microproteins_dashboard.py at commit f742876, no license · at the source

Overview

Authors: Brendan Miller1, Eduardo Vieira de Souza1, Calvin Lau1, Joan M. Vaughan1, Victor J. Pai1, Servando Giraldez2, Andréa Rocha1, Jolene K. Diedrich3, Clodagh C. O’Shea2, David A. Bennett4, Alan Saghatelian1
  1. Clayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies,La Jolla, CA USA
  2. Molecular and Cell Biology Laboratory, The Salk Institute for Biological Studies,La Jolla, CA USA
  3. Department of Integrative Structural and Computational Biology, The Scripps Research Institute,La Jolla, CA USA
  4. Rush Alzheimer’s Disease Center, Rush University Medical Center,Chicago, IL USA
Institutions: Salk Institute for Biological Studies (United States); Scripps Research Institute (United States); Rush University Medical Center (United States)
Journal: Nature aging, volume 6, issue 9, pages 1979-1997
Dates: received 3 October 2025; accepted 28 July 2026; published online 14 September 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s43587-026-01207-x · PMID 42736449 · PMCID PMC13577902 · OpenAlex W7212863263
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), other (modality), human (organism), Alzheimer's / dementia (population)
Methods: Spectral & time-frequency, Statistics, Smoothing, state filtering, decompositions, Preprocessing, Connectivity, fMRI & imaging
Keywords: Ageing, Alzheimer's disease, Gene expression
MeSH: Alzheimer Disease*, Frontal Lobe*, Proteome*, Humans, Mass Spectrometry, Micropeptides, Open Reading Frames, Proteomics (* major topic)
Topic: Alzheimer's disease research and treatments (Physiology, Medicine), according to OpenAlex
Funding: NIGMS (R01GM102491)
Citations: not cited yet (Europe PMC); 59 references in the paper

Abstract

Understanding the molecular basis of neurodegeneration requires a comprehensive map of the genome’s protein-coding output. Although thousands of small open reading frames (ORFs) are translated in the human brain, proteomic evidence for their encoded microproteins (MPs) (≤150 amino acids (aa)) remains limited. Here, we present a brain MP atlas that integrates transcriptomics, mass spectrometry and deep-learning-predicted spectra across more than 600 postmortem frontal cortex samples with and without Alzheimer’s disease (AD). We identified 1,067 MPs absent from reviewed UniProtKB entries with high-confidence spectral support. A subset is differentially expressed in AD independently of the annotated main ORF at the same locus. A small ORF expressed by MKKS encodes a 63-amino-acid MP that is the locus’s predominant translation product and is downregulated in AD; its loss impairs microglial mitochondrial respiration, implicating it in microglial bioenergetics. This atlas expands the annotated brain proteome and provides a resource for studying MPs in aging and neurodegeneration.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 29 matches between paragraphs and lines of code.

huggingface.co/spaces/brmiller/brain-microprotein-atlas-app

License: MIT
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: 5a558dce1875fa365360e0b8329d9c8bb7e2a952, 25 September 2026
Languages: Python (40), Shell (13), R (12)
Size: 256 files, 65 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (Dockerfile, environment.yml, requirements.txt, Results/requirements.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers

brendan-miller-salk/brain-microprotein-atlas

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: f7428766396fb1108ebc2211a512cfed3c0ad259, 25 September 2026
Languages: Python (37), Shell (12), R (11)
Size: 192 files, 60 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (Dockerfile, environment.yml, requirements.txt, Results/requirements.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: pandas (24 files), tidyverse (10 files), NumPy (7 files), ggplot2 (6 files), DESeq2 (5 files), Matplotlib (3 files), SciPy (3 files), BEDTools (2 files), Biopython (2 files), limma (2 files), patchwork (2 files), Plotly (2 files), broom (1 file), pheatmap (1 file), psych (1 file), reshape2 (1 file), scikit-image (1 file), tifffile (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
61 files

Code availability

Code is available on GitHub at https://github.com/brendan-miller-salk/brain-microprotein-atlas. The repository includes analysis scripts, custom pipelines, results, BED/GTF/FASTA files and tools for MP annotation and visualization. The code and data can be further visualized and browsed at https://huggingface.co/spaces/brmiller/brain-microprotein-atlas-app.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 60 scripts, each with its path and the digest of its content;
  • 29 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The data used in this study were in part obtained through the AD Knowledge Portal (https://doi.org/10.7303/9618238). Transcriptomics and associative data were provided by the Rush Alzheimer’s Disease Center, Rush University Medical Center; additional phenotypic data can be requested from www.radc.rush.edu. Transcriptomics data were also generated from postmortem brain tissue collected through the Mount Sinai VA Medical Center Brain Bank and provided by E. Schadt from Mount Sinai School of Medicine. The TMT proteomics data were based on samples provided by the Rush Alzheimer’s Disease Center, Rush University Medical Center. The raw and processed data used in this study are available through multiple public repositories. Data from the AMP-AD Knowledge Portal (https://adknowledgeportal.synapse.org), including the raw and processed datasets from ROSMAP (Synapse ID: syn3219045 (http://www.synapse.org/Synapse:syn3219045)) and MSBB (Synapse ID: syn3159438 (http://www.synapse.org/Synapse:syn3159438)), and methods developed by the AMP-AD Target Discovery Program (supported by the NIA), are shared without publication embargo and are available for secondary analysis; access requires registration and adherence to data use and attribution guidelines (https://adknowledgeportal.synapse.org/#/DataAccess/Instructions). ROSMAP resources can also be requested directly through the Rush Alzheimer’s Disease Center (www.radc.rush.edu). All raw long-read sequencing data generated from 12 postmortem human brain samples are available through Synapse (ID: syn52047893 (http://www.synapse.org/Synapse:syn52047893/wiki/622953)) under restricted access, requiring a free Synapse account. The Ribo-seq data from Duffy et al.5 are available through dbGaP under accession no. phs002489.v1.p1 (https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs002489.v1.p1). The UniProtKB human proteome database (Swiss-Prot and TrEMBL reference sequences, UP000005640 (http://www.uniprot.org/proteomes/UP000005640), downloaded on 20 January 2025) used for the spectral searches is available at https://ftp.uniprot.org/pub/databases/uniprot/previous_releases/. Raw MS data for the HMC3 studies and the UCSD DIA-MS cohort have been deposited at the PRIDE Archive under accession no. PXD071500 (http://proteomecentral.proteomexchange.org/cgi/GetDataset?ID=PXD071500).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 3, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 11 authors, 3 keywords, 8 MeSH terms, 1 funder, 59 references.

Cite

This paper

Miller, B., Vieira de Souza, E., Lau, C., Vaughan, J. M., Pai, V. J., Giraldez, S., Rocha, A., Diedrich, J. K., O’Shea, C. C., Bennett, D. A., & Saghatelian, A. (2026). A microprotein atlas of the human frontal cortex in Alzheimer's disease. Nature aging, 6(9), 1979-1997. https://doi.org/10.1038/s43587-026-01207-x

BibTeX

@article{miller2026microprotein,
author = {Miller, Brendan and Vieira de Souza, Eduardo and Lau, Calvin and Vaughan, Joan M. and Pai, Victor J. and Giraldez, Servando and Rocha, Andréa and Diedrich, Jolene K. and O’Shea, Clodagh C. and Bennett, David A. and Saghatelian, Alan},
title = {{A microprotein atlas of the human frontal cortex in Alzheimer's disease}},
journal = {Nature aging},
year = {2026},
month = sep,
volume = {6},
number = {9},
pages = {1979--1997},
publisher = {Nature Portfolio},
issn = {2662-8465},
doi = {10.1038/s43587-026-01207-x},
url = {https://doi.org/10.1038/s43587-026-01207-x},
pmid = {42736449},
pmcid = {PMC13577902}
}

RIS

TY - JOUR
AU - Miller, Brendan
AU - Vieira de Souza, Eduardo
AU - Lau, Calvin
AU - Vaughan, Joan M.
AU - Pai, Victor J.
AU - Giraldez, Servando
AU - Rocha, Andréa
AU - Diedrich, Jolene K.
AU - O’Shea, Clodagh C.
AU - Bennett, David A.
AU - Saghatelian, Alan
TI - A microprotein atlas of the human frontal cortex in Alzheimer's disease
T2 - Nature aging
J2 - Nat Aging
PY - 2026
DA - 2026/09/14
VL - 6
IS - 9
SP - 1979
EP - 1997
SN - 2662-8465
PB - Nature Portfolio
DO - 10.1038/s43587-026-01207-x
UR - https://doi.org/10.1038/s43587-026-01207-x
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s43587-026-01207-x",
"type": "article-journal",
"title": "A microprotein atlas of the human frontal cortex in Alzheimer's disease",
"container-title": "Nature aging",
"author": [
{
"family": "Miller",
"given": "Brendan"
},
{
"family": "Vieira de Souza",
"given": "Eduardo"
},
{
"family": "Lau",
"given": "Calvin"
},
{
"family": "Vaughan",
"given": "Joan M."
},
{
"family": "Pai",
"given": "Victor J."
},
{
"family": "Giraldez",
"given": "Servando"
},
{
"family": "Rocha",
"given": "Andréa"
},
{
"family": "Diedrich",
"given": "Jolene K."
},
{
"family": "O’Shea",
"given": "Clodagh C."
},
{
"family": "Bennett",
"given": "David A."
},
{
"family": "Saghatelian",
"given": "Alan"
}
],
"container-title-short": "Nat Aging",
"volume": "6",
"issue": "9",
"page": "1979-1997",
"DOI": "10.1038/s43587-026-01207-x",
"PMID": "42736449",
"PMCID": "PMC13577902",
"ISSN": "2662-8465",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s43587-026-01207-x",
"language": "en",
"issued": {
"date-parts": [
[
2026,
9,
14
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Biopython, limma, broom, 12 other tools
[2] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: Biopython, BEDTools, limma, 10 other tools, genetics / omics, 1 reference
[3] doi:10.1038/s41467-026-75571-y [code]
Ribo-ITP enables identification of translons from limited input samples.
Journal: Nature communications
In common: ggplot2, tidyverse, pandas, 2 other tools, genetics / omics, 8 references
[4] doi:10.1073/pnas.2609132123 [code]
A human lysosomal storage disorder toolkit for decoding proteome landscapes in cortical-like and dopaminergic-like induced neurons.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: tifffile, limma, broom, 11 other tools, genetics / omics
[5] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: tifffile, limma, DESeq2, 11 other tools, genetics / omics
[6] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: limma, broom, DESeq2, 10 other tools, genetics / omics, 1 reference
[7] doi:10.1186/s11689-026-09713-0 [code]
DRP1 mutations associated with EMPF1 encephalopathy perturb the transcriptional profile and maturation of cortical neurons.
Journal: Journal of neurodevelopmental disorders
In common: tifffile, limma, DESeq2, 11 other tools
[8] doi:10.1038/s44318-026-00818-9 [code]
FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.
Journal: The EMBO journal
In common: Biopython, BEDTools, limma, 10 other tools, Alzheimer's / dementia
[9] doi:10.1126/sciadv.aed2952 [code]
Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.
Journal: Science advances
In common: Biopython, BEDTools, limma, 10 other tools
[10] doi:10.1002/alz.71558 [code]
Allele specific expression in Alzheimer's disease.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: reshape2, ggplot2, pandas, 2 other tools, synapse.org/synapse:syn3159438, synapse.org/synapse:syn3219045, Alzheimer's / dementia, genetics / omics, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.