OSCR

Connectomic mapping of pharyngeal and gut sensory circuits in adult <i>Drosophila</i>.

Code ↔ Paper

14 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 14 matches
  1. [1] § STAR★Methods › Method details › Clustering sensory axons based on their output connectivity ↔ analyses/3_Coconatfly_aPhN_Cosine.ipynb, lines 1–74 · score 0.79 · Ward D2, hierarchical clustering, FAFB v783, Coconatfly, FlyWire, Notebook
  2. [2] § STAR★Methods › Method details › Clustering sensory axons based on their output connectivity ↔ analyses/2_Coconatfly_PhN_Cosine.ipynb, lines 1–55 · score 0.78 · Ward D2, hierarchical clustering, FAFB v783, FlyWire, Notebook, Coconatfly
  3. [3] § Results › Monosynaptic connections between internal sensory axons and motor or endocrine output neurons ↔ analyses/hops_visualization.ipynb, lines 851–993 · score 0.78 · taste peg, Ir94e, motor neuron, aPhN2, water, bitter
  4. [4] § Results › Monosynaptic connections between internal sensory axons and motor or endocrine output neurons ↔ analyses/hops_visualization.ipynb, lines 851–993 · score 0.75 · taste peg, Ir94e, motor neuron, aPhN1, aPhN2, bitter
  5. [5] § Results › Monosynaptic connections between internal sensory axons and motor or endocrine output neurons ↔ analyses/Full_hops_figure.ipynb, lines 994–1076 · score 0.72 · taste peg, Ir94e, aPhN2, S4, water, bitter
  6. [6] § Results › Monosynaptic connections between internal sensory axons and motor or endocrine output neurons ↔ analyses/Full_hops_figure.ipynb, lines 994–1076 · score 0.67 · taste peg, Ir94e, aPhN1, aPhN2, bitter, sugar
  7. [7] § STAR★Methods › Method details › Identifying putative gut and pharyngeal sensory axons ↔ analyses/1_Coconatfly_aPhN_PhN_Cosine.ipynb, lines 1–55 · score 0.66 · accessory pharyngeal nerve, FlyWire, aPhN, root IDs, sensory neurons, retrieved
  8. [8] § STAR★Methods › Method details › Identifying putative gut and pharyngeal sensory axons ↔ analyses/2_Coconatfly_PhN_Cosine.ipynb, lines 1–55 · score 0.66 · accessory pharyngeal nerve, FlyWire, aPhN, root IDs, retrieved, metadata
  9. [9] § Results › Identification of putative gut and pharyngeal sensory axons in the FAFB connectome ↔ analyses/1_Coconatfly_aPhN_PhN_Cosine.ipynb, lines 1–55 · score 0.64 · accessory pharyngeal nerves, ventral cibarial, dorsal cibarial, labral, rows, LSO
  10. [10] § Results › Identification of putative gut and pharyngeal sensory axons in the FAFB connectome ↔ analyses/5_aPhN_DCSO_connectome_analysis_v2.ipynb, lines 3035–3107 · score 0.59 · SEZ region, aPhN1, aPhN2, PRW, DCSO, neuropil
  11. [11] § STAR★Methods › Method details › Analyzing the minimum number of synaptic hops to motor or endocrine neurons ↔ analyses/Full_hops_figure.ipynb, lines 154–262 · score 0.55 · Sankey diagrams, flow, thickness, hops, node, connections
  12. [12] § STAR★Methods › Method details › Analyzing the minimum number of synaptic hops to motor or endocrine neurons ↔ analyses/5_aPhN_DCSO_connectome_analysis_v2.ipynb, lines 3939–4551 · score 0.53 · root IDs, queries, cell, FlyWire, endocrine, neuron
  13. [13] § Results › Identification of putative gut and pharyngeal sensory axons in the FAFB connectome ↔ analyses/3_Coconatfly_aPhN_Cosine.ipynb, lines 1–74 · score 0.52 · putative DCSO, aPhN1, aPhN2, cosine, LSO, VCSO
  14. [14] § Results › Downstream circuits of the pharyngeal sensory axons ↔ analyses/5_aPhN_DCSO_connectome_analysis_v2.ipynb, lines 3324–3461 · score 0.52 · aPhN1, aPhN2, S2, descending, ascending, DCSO

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 1,080 lines · 41 KB · no license · 3 matches

  1. # %% [markdown]
  2. # ## Workflow overview (from exploratory checks to final portrait figure)
  3. #
  4. # ### Task (what this notebook does)
  5. # This notebook builds **multi-hop synaptic flow summaries** from curated **sensory neuron sets** and visualizes them as **Sankey diagrams**. The earlier cells are **exploratory** (single-panel Sankeys for individual sets and minor variants), while the last section assembles the **final multi-panel portrait figure** (all Sankeys stacked and exported as one SVG for Illustrator).
  6. #
  7. # ---
  8. #
  9. # ### Abbreviations / naming
  10. # - **FlyWire**: the connectome dataset used here.
  11. # - **GRN**: gustatory receptor neuron (input lists loaded from CSVs; each has a `root_id` column).
  12. # - **PhN**: pharyngeal nerve (here: **PhN-SA_v2** sets; 6 CSVs). NOTE: the PhN contains the stomodeal nerve (StN) gut sensory afferents.
  13. # - **PSO**: presumed pharyngeal sensory organ group (here: **aPhN-SA** sets; 3 CSVs labeled `DCSO`, `aPhN1`, `aPhN2`).
  14. # - **MxLbN**: maxillary labellar nerve group (here: 4 GRN lists: `Sugar/Water`, `Bitter`, `Ir94e`, `Taste Peg`).
  15. # - **SA**: sensory axon (used in dataset naming).
  16. # - **`root_id`**: unique neuron identifier (node ID).
  17. # - **`pre_root_id` / `post_root_id`**: presynaptic / postsynaptic neuron IDs in `connections`.
  18. # - **`syn_count`**: number of synapses between a pre→post pair.
  19. # - **Superclass / `super_class`**: broad anatomical/functional category (used to label nodes and color them).
  20. # - **Hop**: one step of directed connectivity (pre→post).
  21. # - **Hop 1**: directly downstream of the input set
  22. # - **Hop 2**: downstream of Hop 1 targets
  23. # - **Hop 3**: downstream of Hop 2 targets
  24. #
  25. # ---
  26. #
  27. # ### Inputs (tables and lists)
  28. # - **Connectome edges**: `flywire_data/connections.csv.(gz|zip)` with columns like `pre_root_id`, `post_root_id`, `syn_count`.
  29. # - **Neuron metadata**: `flywire_data/classification.csv.gz` (subset used: `root_id`, `super_class`).
  30. # - **Seed neuron sets (CSV lists)**:
  31. # - **PhN-SA_v2**: `input/PhN/set_1.csv` … `set_6.csv`
  32. # - **PSO/aPhN-SA**: `input/aPhN-SA/set_1.csv` … `set_3.csv` (mapped to `DCSO`, `aPhN1`, `aPhN2`)
  33. # - **MxLbN-SA**: `input/MxLbN-SA/*.csv` (4 stimulus/modality lists)
  34. #
  35. # ---
  36. #
  37. # ### Core computation (how a Sankey is built)
  38. # 1. **Filter edges by seed set**
  39. # Select rows where `pre_root_id ∈ seed_root_ids`.
  40. #
  41. # 2. **Aggregate to pair-level synapses**
  42. # Group by `(pre_root_id, post_root_id)` and sum `syn_count` so each directed pair has a total weight.
  43. #
  44. # 3. **Threshold weak connections**
  45. # Keep only pairs with **total `syn_count ≥ min_syn`** (default used throughout: `min_syn = 5`).
  46. #
  47. # 4. **Annotate targets with superclass**
  48. # Merge `post_root_id` with `classification` to get `output_super_class`.
  49. #
  50. # 5. **Repeat for downstream hops (1→3)**
  51. # Hop 1 uses the seed IDs; Hop 2 uses Hop 1 target IDs; Hop 3 uses Hop 2 target IDs.
  52. #
  53. # 6. **Summarize flows at the superclass level**
  54. # Convert hop-level edges into flow tables:
  55. # - **Seed → Hop1 superclass totals**
  56. # - **Hop1 superclass → Hop2 superclass totals**
  57. # - **Hop2 superclass → Hop3 superclass totals**
  58. #
  59. # 7. **Plot Sankey with consistent labeling and colors**
  60. # Nodes are labeled by hop, e.g. `1: central`, `2: motor`, `3: endocrine`.
  61. # Colors are assigned by `super_class` so the same superclass has the same color within a figure (and, where implemented, globally).
  62. #
  63. # ---
  64. #
  65. # ### Notebook structure (exploratory → final)
  66. # - **Exploratory section(s)**:
  67. # - Generate **single Sankey diagrams** per set for:
  68. # - the 6 **PhN-SA_v2** sets,
  69. # - the 3 **PSO/aPhN-SA** sets,
  70. # - the 4 **MxLbN-SA** sets.
  71. # - These cells help validate thresholds, hop logic, labels, and layout.
  72. #
  73. # - **Final figure section (for Supplementary Fig. S4)**:
  74. # - Defines `make_sankey_trace(...)` to return a Plotly Sankey trace.
  75. # - Loads **all 13 datasets** (6 PhN + 3 PSO + 4 MxLbN).
  76. # - Uses `plotly.subplots.make_subplots` to stack them into a **13×1 portrait layout**.
  77. # - Exports a **single SVG** (intended for Adobe Illustrator edits).
  78. #
  79. # ---
  80. #
  81. # ### Outputs
  82. # - Individual Sankey panels are shown interactively, and the final portrait figure is written to:
  83. # - `./Giakoumas-et-al/output/figures/fig_S4/all_sankeys_portrait.svg`
  84. # %%
  85. from pathlib import Path
  86. import pandas as pd
  87. import numpy as np
  88. import plotly.graph_objects as go
  89. import plotly.express as px
  90. def find_project_root() -> Path:
  91. current = Path.cwd().resolve()
  92. for candidate in (current, *current.parents):
  93. if (candidate / 'flywire_data').exists() and (candidate / 'input').exists():
  94. return candidate
  95. raise FileNotFoundError('Could not find the Giakoumas-et-al project root from the current working directory.')
  96. PROJECT_ROOT = find_project_root()
  97. DATA_DIR = PROJECT_ROOT / 'flywire_data'
  98. INPUT_DIR = PROJECT_ROOT / 'input'
  99. OUTPUT_DIR = PROJECT_ROOT / 'output' / 'figures' / 'fig_S4'
  100. OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
  101. CONNECTIONS_PATH = DATA_DIR / 'connections.csv'
  102. if not CONNECTIONS_PATH.exists():
  103. CONNECTIONS_PATH = DATA_DIR / 'connections.csv.zip'
  104. CLASSIFICATION_PATH = DATA_DIR / 'classification.csv.gz'
  105. # ----------------------------------------------------------------------------
  106. # 0. Load FlyWire connectome and super_class classification
  107. # ----------------------------------------------------------------------------
  108. connections = pd.read_csv(CONNECTIONS_PATH)
  109. classification = pd.read_csv(CLASSIFICATION_PATH)[['root_id','super_class']]
  110. # ----------------------------------------------------------------------------
  111. # 1. Build a GLOBAL color map for every superclass once, up-front
  112. # ----------------------------------------------------------------------------
  113. all_classes_global = sorted(classification['super_class'].unique())
  114. palette = px.colors.qualitative.Safe
  115. GLOBAL_COLOR_MAP = {
  116. cls: palette[i % len(palette)]
  117. for i, cls in enumerate(all_classes_global)
  118. }
  119. # ----------------------------------------------------------------------------
  120. # 2. Load your six PhN-SA_v2 sets (each CSV has a 'root_id' column)
  121. # ----------------------------------------------------------------------------
  122. sets = {
  123. f'Set{i}': pd.read_csv(INPUT_DIR / 'PhN' / f'set_{i}.csv')
  124. for i in range(1, 7)
  125. }
  126. # ----------------------------------------------------------------------------
  127. # Helper: filter connections by source IDs, threshold syn_count, attach superclass
  128. # ----------------------------------------------------------------------------
  129. def build_hop_df(src_ids, connections, classification, min_syn=5):
  130. df = connections[connections['pre_root_id'].isin(src_ids)]
  131. summed = (
  132. df.groupby(['pre_root_id','post_root_id'], as_index=False)
  133. .agg({'syn_count':'sum'})
  134. .query('syn_count >= @min_syn')
  135. )
  136. merged = pd.merge(
  137. summed,
  138. classification.rename(
  139. columns={'root_id':'post_root_id','super_class':'output_super_class'}
  140. ),
  141. on='post_root_id', how='left'
  142. )
  143. return merged[['pre_root_id','post_root_id','output_super_class','syn_count']]
  144. # ----------------------------------------------------------------------------
  145. # Core: build and show a Sankey diagram with consistent colors
  146. # ----------------------------------------------------------------------------
  147. def plot_sankey_dynamic(grn_df, title, connections, classification, min_syn=5):
  148. # hops
  149. df1 = build_hop_df(grn_df['root_id'], connections, classification, min_syn)
  150. df2 = build_hop_df(df1['post_root_id'].unique(), connections, classification, min_syn)
  151. df3 = build_hop_df(df2['post_root_id'].unique(), connections, classification, min_syn)
  152. # summarize flows
  153. flow1 = (
  154. df1.groupby('output_super_class')['syn_count']
  155. .sum().reset_index(name='count')
  156. .assign(source=title)
  157. )
  158. m12 = pd.merge(df1, df2,
  159. left_on='post_root_id', right_on='pre_root_id',
  160. suffixes=('_1','_2'))
  161. flow2 = (
  162. m12.groupby(['output_super_class_1','output_super_class_2'])['syn_count_2']
  163. .sum().reset_index(name='count')
  164. )
  165. m23 = pd.merge(df2, df3,
  166. left_on='post_root_id', right_on='pre_root_id',
  167. suffixes=('_2','_3'))
  168. flow3 = (
  169. m23.groupby(['output_super_class_2','output_super_class_3'])['syn_count_3']
  170. .sum().reset_index(name='count')
  171. )
  172. # nodes
  173. col1 = [title]
  174. col2 = [f"1: {c}" for c in sorted(df1['output_super_class'].unique())]
  175. col3 = [f"2: {c}" for c in sorted(df2['output_super_class'].unique())]
  176. col4 = [f"3: {c}" for c in sorted(df3['output_super_class'].unique())]
  177. nodes = col1 + col2 + col3 + col4
  178. idx = {node:i for i,node in enumerate(nodes)}
  179. # colors: reuse the GLOBAL_COLOR_MAP
  180. node_colors = [
  181. 'lightgrey' if n == title else GLOBAL_COLOR_MAP[n.split(': ',1)[1]]
  182. for n in nodes
  183. ]
  184. # links
  185. source, target, value, link_colors = [], [], [], []
  186. def add_links(df, src_col, tgt_col):
  187. for _, r in df.iterrows():
  188. s = idx[r[src_col]]
  189. t = idx[r[tgt_col]]
  190. source.append(s)
  191. target.append(t)
  192. value.append(r['count'])
  193. link_colors.append(node_colors[s].replace('rgb','rgba').replace(')',',0.5)'))
  194. # flow1 → 1:class
  195. flow1 = flow1.rename(columns={'source':'src','output_super_class':'dst'})
  196. flow1['dst'] = flow1['dst'].map(lambda c: f"1: {c}")
  197. add_links(flow1, 'src', 'dst')
  198. # flow2 → 2:class
  199. flow2 = flow2.rename(columns={
  200. 'output_super_class_1':'src','output_super_class_2':'dst'
  201. })
  202. flow2['src'] = flow2['src'].map(lambda c: f"1: {c}")
  203. flow2['dst'] = flow2['dst'].map(lambda c: f"2: {c}")
  204. add_links(flow2, 'src', 'dst')
  205. # flow3 → 3:class
  206. flow3 = flow3.rename(columns={
  207. 'output_super_class_2':'src','output_super_class_3':'dst'
  208. })
  209. flow3['src'] = flow3['src'].map(lambda c: f"2: {c}")
  210. flow3['dst'] = flow3['dst'].map(lambda c: f"3: {c}")
  211. add_links(flow3, 'src', 'dst')
  212. # hover
  213. incoming = dict.fromkeys(nodes, 0)
  214. outgoing = dict.fromkeys(nodes, 0)
  215. for s,t,v in zip(source, target, value):
  216. outgoing[nodes[s]] += v
  217. incoming[nodes[t]] += v
  218. customdata = [f"Incoming: {incoming[n]}<br>Outgoing: {outgoing[n]}" for n in nodes]
  219. # positions
  220. x = [0.0]*len(col1) + [0.33]*len(col2) + [0.66]*len(col3) + [1.0]*len(col4)
  221. y = []
  222. for col in (col1, col2, col3, col4):
  223. n = len(col)
  224. y.extend([0.5] if n==1 else list(np.linspace(0,1,n)))
  225. # draw
  226. fig = go.Figure(go.Sankey(
  227. arrangement='snap',
  228. node=dict(
  229. label=nodes,
  230. x=x, y=y,
  231. color=node_colors,
  232. pad=15,
  233. thickness=20,
  234. line=dict(color='black', width=0.5),
  235. customdata=customdata,
  236. hovertemplate='%{customdata}<extra>%{label}</extra>'
  237. ),
  238. link=dict(source=source, target=target, value=value, color=link_colors)
  239. ))
  240. fig.update_layout(title_text=f"{title}", font_size=14)
  241. fig.write_image(OUTPUT_DIR / f"{title.replace('/','_')}.svg", width=1200, height=800, scale=2)
  242. fig.show()
  243. # ----------------------------------------------------------------------------
  244. # 3. Generate a Sankey for each of the six PhN-SA_v2 sets
  245. # ----------------------------------------------------------------------------
  246. for label, df in sets.items():
  247. plot_sankey_dynamic(df, label, connections, classification)
  248. # %% [markdown]
  249. # ### Generate a Sankey for each PSO set
  250. # %%
  251. from pathlib import Path
  252. import pandas as pd
  253. import numpy as np
  254. import plotly.graph_objects as go
  255. import plotly.express as px
  256. if 'find_project_root' not in globals():
  257. def find_project_root() -> Path:
  258. current = Path.cwd().resolve()
  259. for candidate in (current, *current.parents):
  260. if (candidate / 'flywire_data').exists() and (candidate / 'input').exists():
  261. return candidate
  262. raise FileNotFoundError('Could not find the Giakoumas-et-al project root from the current working directory.')
  263. PROJECT_ROOT = find_project_root()
  264. DATA_DIR = PROJECT_ROOT / 'flywire_data'
  265. INPUT_DIR = PROJECT_ROOT / 'input'
  266. OUTPUT_DIR = PROJECT_ROOT / 'output' / 'figures' / 'fig_S4'
  267. OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
  268. CONNECTIONS_PATH = DATA_DIR / 'connections.csv'
  269. if not CONNECTIONS_PATH.exists():
  270. CONNECTIONS_PATH = DATA_DIR / 'connections.csv.zip'
  271. # ----------------------------------------------------------------------------
  272. # 0. Load FlyWire connectome and super_class classification
  273. # ----------------------------------------------------------------------------
  274. connections = pd.read_csv(CONNECTIONS_PATH)
  275. classification = pd.read_csv(DATA_DIR / 'classification.csv.gz')[['root_id','super_class']]
  276. # ----------------------------------------------------------------------------
  277. # 1. Load your three PSO sets (each CSV has a 'root_id' column)
  278. # ----------------------------------------------------------------------------
  279. set_1 = pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_1.csv')
  280. set_2 = pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_2.csv')
  281. set_3 = pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_3.csv')
  282. sets = {
  283. 'DCSO': set_1,
  284. 'aPhN1': set_2,
  285. 'aPhN2': set_3,
  286. }
  287. # ----------------------------------------------------------------------------
  288. # Helper: filter connections by source IDs, threshold syn_count, attach superclass
  289. # ----------------------------------------------------------------------------
  290. def build_hop_df(src_ids, connections, classification, min_syn=5):
  291. df = connections[connections['pre_root_id'].isin(src_ids)]
  292. summed = (
  293. df.groupby(['pre_root_id','post_root_id'], as_index=False)
  294. .agg({'syn_count':'sum'})
  295. .query('syn_count >= @min_syn')
  296. )
  297. merged = pd.merge(
  298. summed,
  299. classification.rename(
  300. columns={'root_id':'post_root_id','super_class':'output_super_class'}
  301. ),
  302. on='post_root_id', how='left'
  303. )
  304. return merged[['pre_root_id','post_root_id','output_super_class','syn_count']]
  305. # ----------------------------------------------------------------------------
  306. # Core: build and show a Sankey diagram with maximal vertical spacing
  307. # ----------------------------------------------------------------------------
  308. def plot_sankey_dynamic(grn_df, title, connections, classification, min_syn=5):
  309. # 1) build hops
  310. df1 = build_hop_df(grn_df['root_id'], connections, classification, min_syn)
  311. df2 = build_hop_df(df1['post_root_id'].unique(), connections, classification, min_syn)
  312. df3 = build_hop_df(df2['post_root_id'].unique(), connections, classification, min_syn)
  313. # 2) summarize flows
  314. flow1 = (
  315. df1.groupby('output_super_class')['syn_count']
  316. .sum().reset_index(name='count')
  317. .assign(source=title)
  318. )
  319. m12 = pd.merge(df1, df2, left_on='post_root_id', right_on='pre_root_id', suffixes=('_1','_2'))
  320. flow2 = (
  321. m12.groupby(['output_super_class_1','output_super_class_2'])['syn_count_2']
  322. .sum().reset_index(name='count')
  323. )
  324. m23 = pd.merge(df2, df3, left_on='post_root_id', right_on='pre_root_id', suffixes=('_2','_3'))
  325. flow3 = (
  326. m23.groupby(['output_super_class_2','output_super_class_3'])['syn_count_3']
  327. .sum().reset_index(name='count')
  328. )
  329. # 3) build node labels
  330. col1 = [title]
  331. col2 = [f"1: {c}" for c in sorted(df1['output_super_class'].unique())]
  332. col3 = [f"2: {c}" for c in sorted(df2['output_super_class'].unique())]
  333. col4 = [f"3: {c}" for c in sorted(df3['output_super_class'].unique())]
  334. nodes = col1 + col2 + col3 + col4
  335. idx = {n:i for i,n in enumerate(nodes)}
  336. # 4) color mapping
  337. palette = px.colors.qualitative.Safe
  338. all_classes = sorted(
  339. set(df1['output_super_class']) |
  340. set(df2['output_super_class']) |
  341. set(df3['output_super_class'])
  342. )
  343. color_map = {cls: palette[i % len(palette)] for i,cls in enumerate(all_classes)}
  344. node_colors = ['lightgrey' if n==title else color_map[n.split(': ',1)[1]] for n in nodes]
  345. # 5) assemble links
  346. source, target, value, link_colors = [], [], [], []
  347. def add_links(flow_df, src_col, tgt_col):
  348. for _, r in flow_df.iterrows():
  349. s = idx[r[src_col]]
  350. t = idx[r[tgt_col]]
  351. source.append(s)
  352. target.append(t)
  353. value.append(r['count'])
  354. link_colors.append(node_colors[s].replace('rgb','rgba').replace(')',',0.5)'))
  355. flow1 = flow1.rename(columns={'source':'src','output_super_class':'dst'})
  356. flow1['dst'] = flow1['dst'].map(lambda c: f"1: {c}")
  357. add_links(flow1, 'src','dst')
  358. flow2 = flow2.rename(columns={'output_super_class_1':'src','output_super_class_2':'dst'})
  359. flow2['src'] = flow2['src'].map(lambda c: f"1: {c}")
  360. flow2['dst'] = flow2['dst'].map(lambda c: f"2: {c}")
  361. add_links(flow2, 'src','dst')
  362. flow3 = flow3.rename(columns={'output_super_class_2':'src','output_super_class_3':'dst'})
  363. flow3['src'] = flow3['src'].map(lambda c: f"2: {c}")
  364. flow3['dst'] = flow3['dst'].map(lambda c: f"3: {c}")
  365. add_links(flow3, 'src','dst')
  366. # 6) hover info
  367. incoming = dict.fromkeys(nodes,0)
  368. outgoing = dict.fromkeys(nodes,0)
  369. for s,t,v in zip(source,target,value):
  370. outgoing[nodes[s]] += v
  371. incoming[nodes[t]] += v
  372. customdata = [f"Incoming: {incoming[n]}<br>Outgoing: {outgoing[n]}" for n in nodes]
  373. # 7) x positions
  374. x = [0.0]*len(col1) + [0.33]*len(col2) + [0.66]*len(col3) + [1.0]*len(col4)
  375. # 8) y positions: spread evenly down the column
  376. y = []
  377. for col in (col1, col2, col3, col4):
  378. n = len(col)
  379. if n == 1:
  380. y.append(0.5)
  381. else:
  382. y.extend(np.linspace(0, 1, n))
  383. # 9) draw Sankey with freeform arrangement
  384. fig = go.Figure(go.Sankey(
  385. arrangement='snap',
  386. node=dict(
  387. label=nodes,
  388. x=x, y=y,
  389. color=node_colors,
  390. pad=15, thickness=20,
  391. line=dict(color='black', width=0.5),
  392. customdata=customdata,
  393. hovertemplate='%{customdata}<extra>%{label}</extra>'
  394. ),
  395. link=dict(source=source, target=target, value=value, color=link_colors)
  396. ))
  397. fig.update_layout(title_text=f"{title}", font_size=14)
  398. fig.write_image(OUTPUT_DIR / f"{title.replace('/','_')}.svg", width=1200, height=800, scale=2)
  399. fig.show()
  400. # ----------------------------------------------------------------------------
  401. # 10. Generate a Sankey for each PSO set
  402. # ----------------------------------------------------------------------------
  403. for label, df in sets.items():
  404. plot_sankey_dynamic(df, label, connections, classification)
  405. # %% [markdown]
  406. # ### Generate Sankey for each MxLbN set
  407. # %%
  408. import pandas as pd
  409. import plotly.graph_objects as go
  410. from pathlib import Path
  411. import plotly.express as px
  412. if 'find_project_root' not in globals():
  413. def find_project_root() -> Path:
  414. current = Path.cwd().resolve()
  415. for candidate in (current, *current.parents):
  416. if (candidate / 'flywire_data').exists() and (candidate / 'input').exists():
  417. return candidate
  418. raise FileNotFoundError('Could not find the Giakoumas-et-al project root from the current working directory.')
  419. PROJECT_ROOT = find_project_root()
  420. DATA_DIR = PROJECT_ROOT / 'flywire_data'
  421. INPUT_DIR = PROJECT_ROOT / 'input'
  422. CONNECTIONS_PATH = DATA_DIR / 'connections.csv'
  423. if not CONNECTIONS_PATH.exists():
  424. CONNECTIONS_PATH = DATA_DIR / 'connections.csv.zip'
  425. # ----------------------------------------------------------------------------
  426. # 0. Load FlyWire connectome and classification
  427. # ----------------------------------------------------------------------------
  428. connections = pd.read_csv(CONNECTIONS_PATH)
  429. classification = pd.read_csv(DATA_DIR / 'classification.csv.gz')[['root_id','super_class']]
  430. # ----------------------------------------------------------------------------
  431. # 1. Load your four MxLbN GRN lists (each CSV has a 'root_id' column).
  432. # ----------------------------------------------------------------------------
  433. sugar_water = pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'sugar_water_GRNs.csv')
  434. bitter = pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'bitter_GRNs.csv')
  435. ir94e = pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'Ir94e_GRNs.csv')
  436. taste_peg = pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'taste_peg_GRNs.csv')
  437. sets = {
  438. 'Sugar/Water': sugar_water,
  439. 'Bitter': bitter,
  440. 'Ir94e': ir94e,
  441. 'Taste Peg': taste_peg,
  442. }
  443. # ----------------------------------------------------------------------------
  444. # Helper: filter connections by source IDs, threshold syn_count, attach superclass
  445. # ----------------------------------------------------------------------------
  446. def build_hop_df(src_ids, connections, classification, min_syn=5):
  447. df = connections[connections['pre_root_id'].isin(src_ids)]
  448. summed = (
  449. df.groupby(['pre_root_id','post_root_id'], as_index=False)
  450. .agg({'syn_count':'sum'})
  451. .query('syn_count >= @min_syn')
  452. )
  453. merged = pd.merge(
  454. summed,
  455. classification.rename(columns={'root_id':'post_root_id','super_class':'output_super_class'}),
  456. on='post_root_id', how='left'
  457. )
  458. return merged[['pre_root_id','post_root_id','output_super_class','syn_count']]
  459. # ----------------------------------------------------------------------------
  460. # Core: build and show a Sankey diagram using Plotly's Safe palette
  461. # ----------------------------------------------------------------------------
  462. def plot_sankey_dynamic(grn_df, title, connections, classification, min_syn=5):
  463. # build hop-level dataframes
  464. df1 = build_hop_df(grn_df['root_id'], connections, classification, min_syn)
  465. df2 = build_hop_df(df1['post_root_id'].unique(), connections, classification, min_syn)
  466. df3 = build_hop_df(df2['post_root_id'].unique(), connections, classification, min_syn)
  467. # summarize flows
  468. flow1 = (
  469. df1.groupby('output_super_class')['syn_count']
  470. .sum()
  471. .reset_index(name='count')
  472. .assign(source=title)
  473. )
  474. m12 = pd.merge(
  475. df1, df2,
  476. left_on='post_root_id', right_on='pre_root_id',
  477. suffixes=('_1','_2')
  478. )
  479. flow2 = (
  480. m12.groupby(['output_super_class_1','output_super_class_2'])['syn_count_2']
  481. .sum()
  482. .reset_index(name='count')
  483. )
  484. m23 = pd.merge(
  485. df2, df3,
  486. left_on='post_root_id', right_on='pre_root_id',
  487. suffixes=('_2','_3')
  488. )
  489. flow3 = (
  490. m23.groupby(['output_super_class_2','output_super_class_3'])['syn_count_3']
  491. .sum()
  492. .reset_index(name='count')
  493. )
  494. # assemble nodes
  495. col1 = [title]
  496. col2 = [f"1: {c}" for c in sorted(df1['output_super_class'].unique())]
  497. col3 = [f"2: {c}" for c in sorted(df2['output_super_class'].unique())]
  498. col4 = [f"3: {c}" for c in sorted(df3['output_super_class'].unique())]
  499. nodes = col1 + col2 + col3 + col4
  500. idx = {node:i for i,node in enumerate(nodes)}
  501. # assign colors via Safe palette
  502. palette = px.colors.qualitative.Safe
  503. all_classes = sorted(
  504. set(df1['output_super_class'])
  505. | set(df2['output_super_class'])
  506. | set(df3['output_super_class'])
  507. )
  508. color_map = {cls: palette[i % len(palette)] for i,cls in enumerate(all_classes)}
  509. node_colors = [
  510. 'lightgrey' if n == title else color_map[n.split(': ',1)[1]]
  511. for n in nodes
  512. ]
  513. # build link lists
  514. source, target, value, link_colors = [], [], [], []
  515. # flow1 links
  516. for _, r in flow1.iterrows():
  517. s = idx[title]
  518. t = idx[f"1: {r['output_super_class']}"]
  519. source.append(s); target.append(t); value.append(r['count'])
  520. link_colors.append(node_colors[s].replace('rgb','rgba').replace(')',',0.5)'))
  521. # flow2 links
  522. for _, r in flow2.iterrows():
  523. s = idx[f"1: {r['output_super_class_1']}"]
  524. t = idx[f"2: {r['output_super_class_2']}"]
  525. source.append(s); target.append(t); value.append(r['count'])
  526. link_colors.append(node_colors[s].replace('rgb','rgba').replace(')',',0.5)'))
  527. # flow3 links
  528. for _, r in flow3.iterrows():
  529. s = idx[f"2: {r['output_super_class_2']}"]
  530. t = idx[f"3: {r['output_super_class_3']}"]
  531. source.append(s); target.append(t); value.append(r['count'])
  532. link_colors.append(node_colors[s].replace('rgb','rgba').replace(')',',0.5)'))
  533. # compute hover info
  534. incoming, outgoing = dict.fromkeys(nodes,0), dict.fromkeys(nodes,0)
  535. for s,t,v in zip(source,target,value):
  536. outgoing[nodes[s]] += v
  537. incoming[nodes[t]] += v
  538. customdata = [
  539. f"Incoming: {incoming[n]}<br>Outgoing: {outgoing[n]}"
  540. for n in nodes
  541. ]
  542. # set x-positions
  543. x = [0.0]*len(col1) + [0.33]*len(col2) + [0.66]*len(col3) + [1.0]*len(col4)
  544. # plot Sankey
  545. fig = go.Figure(go.Sankey(
  546. arrangement='snap',
  547. node=dict(
  548. label=nodes,
  549. x=x,
  550. color=node_colors,
  551. pad=15,
  552. thickness=20,
  553. line=dict(color='black', width=0.5),
  554. customdata=customdata,
  555. hovertemplate='%{customdata}<extra>%{label}</extra>'
  556. ),
  557. link=dict(source=source, target=target, value=value, color=link_colors)
  558. ))
  559. fig.update_layout(title_text=title, font_size=14)
  560. fig.show()
  561. # ----------------------------------------------------------------------------
  562. # Generate Sankey for each MxLbN set
  563. # ----------------------------------------------------------------------------
  564. datasets = [
  565. ('Sugar/Water', sugar_water),
  566. ('Bitter', bitter),
  567. ('Ir94e', ir94e),
  568. ('Taste Peg', taste_peg),
  569. ]
  570. for label, df in datasets:
  571. plot_sankey_dynamic(df, label, connections, classification)
  572. # %% [markdown]
  573. # ### Generate Sankey for each PSOS set
  574. # %%
  575. from pathlib import Path
  576. import pandas as pd
  577. import numpy as np
  578. import plotly.graph_objects as go
  579. import plotly.express as px
  580. from plotly.subplots import make_subplots
  581. if 'find_project_root' not in globals():
  582. def find_project_root() -> Path:
  583. current = Path.cwd().resolve()
  584. for candidate in (current, *current.parents):
  585. if (candidate / 'flywire_data').exists() and (candidate / 'input').exists():
  586. return candidate
  587. raise FileNotFoundError('Could not find the Giakoumas-et-al project root from the current working directory.')
  588. PROJECT_ROOT = find_project_root()
  589. DATA_DIR = PROJECT_ROOT / 'flywire_data'
  590. INPUT_DIR = PROJECT_ROOT / 'input'
  591. CONNECTIONS_PATH = DATA_DIR / 'connections.csv'
  592. if not CONNECTIONS_PATH.exists():
  593. CONNECTIONS_PATH = DATA_DIR / 'connections.csv.zip'
  594. if 'connections' not in globals():
  595. connections = pd.read_csv(CONNECTIONS_PATH)
  596. if 'classification' not in globals():
  597. classification = pd.read_csv(DATA_DIR / 'classification.csv.gz')[['root_id','super_class']]
  598. if 'make_sankey_trace' not in globals():
  599. def make_sankey_trace(grn_df, title, connections, classification, min_syn=5):
  600. def build_hop_df(src_ids):
  601. df = connections[connections['pre_root_id'].isin(src_ids)]
  602. summed = (
  603. df.groupby(['pre_root_id', 'post_root_id'], as_index=False)
  604. .agg({'syn_count': 'sum'})
  605. )
  606. summed = summed[summed['syn_count'] >= min_syn]
  607. merged = pd.merge(
  608. summed,
  609. classification.rename(
  610. columns={'root_id': 'post_root_id', 'super_class': 'output_super_class'}
  611. ),
  612. on='post_root_id', how='left'
  613. )
  614. return merged[['pre_root_id', 'post_root_id', 'output_super_class', 'syn_count']]
  615. df1 = build_hop_df(grn_df['root_id'])
  616. df2 = build_hop_df(df1['post_root_id'].unique())
  617. df3 = build_hop_df(df2['post_root_id'].unique())
  618. flow1 = (
  619. df1.groupby('output_super_class')['syn_count']
  620. .sum().reset_index(name='count')
  621. .assign(source=title)
  622. )
  623. m12 = pd.merge(
  624. df1, df2,
  625. left_on='post_root_id', right_on='pre_root_id',
  626. suffixes=('_1', '_2')
  627. )
  628. flow2 = (
  629. m12.groupby(['output_super_class_1', 'output_super_class_2'])['syn_count_2']
  630. .sum().reset_index(name='count')
  631. )
  632. m23 = pd.merge(
  633. df2, df3,
  634. left_on='post_root_id', right_on='pre_root_id',
  635. suffixes=('_2', '_3')
  636. )
  637. flow3 = (
  638. m23.groupby(['output_super_class_2', 'output_super_class_3'])['syn_count_3']
  639. .sum().reset_index(name='count')
  640. )
  641. col1 = [title]
  642. col2 = [f"1: {c}" for c in sorted(df1['output_super_class'].unique())]
  643. col3 = [f"2: {c}" for c in sorted(df2['output_super_class'].unique())]
  644. col4 = [f"3: {c}" for c in sorted(df3['output_super_class'].unique())]
  645. nodes = col1 + col2 + col3 + col4
  646. idx = {n: i for i, n in enumerate(nodes)}
  647. palette = px.colors.qualitative.Safe
  648. all_classes = sorted(
  649. set(df1['output_super_class']) |
  650. set(df2['output_super_class']) |
  651. set(df3['output_super_class'])
  652. )
  653. color_map = {cls: palette[i % len(palette)] for i, cls in enumerate(all_classes)}
  654. node_colors = [
  655. 'lightgrey' if n == title else color_map[n.split(': ', 1)[1]]
  656. for n in nodes
  657. ]
  658. source, target, value, link_colors = [], [], [], []
  659. flow1 = flow1.rename(columns={'source': 'src', 'output_super_class': 'dst'})
  660. flow1['dst'] = flow1['dst'].map(lambda c: f"1: {c}")
  661. for _, r in flow1.iterrows():
  662. s = idx[r['src']]
  663. t = idx[r['dst']]
  664. source.append(s)
  665. target.append(t)
  666. value.append(r['count'])
  667. link_colors.append(node_colors[s].replace('rgb', 'rgba').replace(')', ',0.5)'))
  668. flow2 = flow2.rename(columns={'output_super_class_1': 'src', 'output_super_class_2': 'dst'})
  669. flow2['src'] = flow2['src'].map(lambda c: f"1: {c}")
  670. flow2['dst'] = flow2['dst'].map(lambda c: f"2: {c}")
  671. for _, r in flow2.iterrows():
  672. s = idx[r['src']]
  673. t = idx[r['dst']]
  674. source.append(s)
  675. target.append(t)
  676. value.append(r['count'])
  677. link_colors.append(node_colors[s].replace('rgb', 'rgba').replace(')', ',0.5)'))
  678. flow3 = flow3.rename(columns={'output_super_class_2': 'src', 'output_super_class_3': 'dst'})
  679. flow3['src'] = flow3['src'].map(lambda c: f"2: {c}")
  680. flow3['dst'] = flow3['dst'].map(lambda c: f"3: {c}")
  681. for _, r in flow3.iterrows():
  682. s = idx[r['src']]
  683. t = idx[r['dst']]
  684. source.append(s)
  685. target.append(t)
  686. value.append(r['count'])
  687. link_colors.append(node_colors[s].replace('rgb', 'rgba').replace(')', ',0.5)'))
  688. incoming = dict.fromkeys(nodes, 0)
  689. outgoing = dict.fromkeys(nodes, 0)
  690. for s_, t_, v_ in zip(source, target, value):
  691. outgoing[nodes[s_]] += v_
  692. incoming[nodes[t_]] += v_
  693. customdata = [f"Incoming: {incoming[n]}<br>Outgoing: {outgoing[n]}" for n in nodes]
  694. x = [0.0] * len(col1) + [0.33] * len(col2) + [0.66] * len(col3) + [1.0] * len(col4)
  695. y = []
  696. for col in (col1, col2, col3, col4):
  697. n = len(col)
  698. if n == 1:
  699. y.append(0.5)
  700. else:
  701. y.extend(list(np.linspace(0, 1, n)))
  702. return go.Sankey(
  703. name=title,
  704. arrangement='snap',
  705. node=dict(
  706. label=nodes,
  707. x=x,
  708. y=y,
  709. color=node_colors,
  710. pad=8,
  711. thickness=12,
  712. line=dict(color='black', width=0.3),
  713. customdata=customdata,
  714. hovertemplate='%{customdata}<extra>%{label}</extra>'
  715. ),
  716. link=dict(source=source, target=target, value=value, color=link_colors)
  717. )
  718. # Only PSOs:
  719. psos = [
  720. ('DCSO', pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_1.csv')),
  721. ('aPhN1', pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_2.csv')),
  722. ('aPhN2', pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_3.csv')),
  723. ]
  724. fig = make_subplots(
  725. rows=3, cols=1,
  726. vertical_spacing=0.02,
  727. specs=[[{"type": "sankey"}]] * 3
  728. )
  729. for i, (label, df) in enumerate(psos, start=1):
  730. sankey_trace = make_sankey_trace(
  731. grn_df=df,
  732. title=label,
  733. connections=connections,
  734. classification=classification,
  735. min_syn=5,
  736. )
  737. fig.add_trace(sankey_trace, row=i, col=1)
  738. fig.update_layout(
  739. height=800,
  740. width=600,
  741. margin=dict(l=20, r=20, t=30, b=20),
  742. font=dict(size=10),
  743. title="PSO Sets Sankeys (stacked)",
  744. )
  745. fig.show()
  746. # %% [markdown]
  747. # ### Main figure for Supplementary Figure S4 to be edited in Adobe Illustrator: all 13 Sankey diagrams in one portrait figure
  748. # %%
  749. import pandas as pd
  750. import numpy as np
  751. import plotly.graph_objects as go
  752. from plotly.subplots import make_subplots
  753. import plotly.express as px
  754. def make_sankey_trace(grn_df, title, connections, classification, min_syn=5):
  755. # Build hop‐level dataframes without using .query()
  756. def build_hop_df(src_ids):
  757. df = connections[connections['pre_root_id'].isin(src_ids)]
  758. summed = (
  759. df.groupby(['pre_root_id', 'post_root_id'], as_index=False)
  760. .agg({'syn_count': 'sum'})
  761. )
  762. summed = summed[summed['syn_count'] >= min_syn]
  763. merged = pd.merge(
  764. summed,
  765. classification.rename(
  766. columns={'root_id': 'post_root_id', 'super_class': 'output_super_class'}
  767. ),
  768. on='post_root_id', how='left'
  769. )
  770. return merged[['pre_root_id', 'post_root_id', 'output_super_class', 'syn_count']]
  771. # 1) Build the three hops
  772. df1 = build_hop_df(grn_df['root_id'])
  773. df2 = build_hop_df(df1['post_root_id'].unique())
  774. df3 = build_hop_df(df2['post_root_id'].unique())
  775. # 2) Summarize flows
  776. flow1 = (
  777. df1.groupby('output_super_class')['syn_count']
  778. .sum().reset_index(name='count')
  779. .assign(source=title)
  780. )
  781. m12 = pd.merge(
  782. df1, df2,
  783. left_on='post_root_id', right_on='pre_root_id',
  784. suffixes=('_1', '_2')
  785. )
  786. flow2 = (
  787. m12.groupby(['output_super_class_1', 'output_super_class_2'])['syn_count_2']
  788. .sum().reset_index(name='count')
  789. )
  790. m23 = pd.merge(
  791. df2, df3,
  792. left_on='post_root_id', right_on='pre_root_id',
  793. suffixes=('_2', '_3')
  794. )
  795. flow3 = (
  796. m23.groupby(['output_super_class_2', 'output_super_class_3'])['syn_count_3']
  797. .sum().reset_index(name='count')
  798. )
  799. # 3) Node labels
  800. col1 = [title]
  801. col2 = [f"1: {c}" for c in sorted(df1['output_super_class'].unique())]
  802. col3 = [f"2: {c}" for c in sorted(df2['output_super_class'].unique())]
  803. col4 = [f"3: {c}" for c in sorted(df3['output_super_class'].unique())]
  804. nodes = col1 + col2 + col3 + col4
  805. idx = {n: i for i, n in enumerate(nodes)}
  806. # 4) Node colors
  807. palette = px.colors.qualitative.Safe
  808. all_classes = sorted(
  809. set(df1['output_super_class']) |
  810. set(df2['output_super_class']) |
  811. set(df3['output_super_class'])
  812. )
  813. color_map = {cls: palette[i % len(palette)] for i, cls in enumerate(all_classes)}
  814. node_colors = [
  815. 'lightgrey' if n == title else color_map[n.split(': ', 1)[1]]
  816. for n in nodes
  817. ]
  818. # 5) Build link lists
  819. source, target, value, link_colors = [], [], [], []
  820. # flow1 → “1:Class”
  821. flow1 = flow1.rename(columns={'source': 'src', 'output_super_class': 'dst'})
  822. flow1['dst'] = flow1['dst'].map(lambda c: f"1: {c}")
  823. for _, r in flow1.iterrows():
  824. s = idx[r['src']]
  825. t = idx[r['dst']]
  826. source.append(s)
  827. target.append(t)
  828. value.append(r['count'])
  829. link_colors.append(node_colors[s].replace('rgb', 'rgba').replace(')', ',0.5)'))
  830. # flow2 → “2:Class”
  831. flow2 = flow2.rename(columns={'output_super_class_1': 'src', 'output_super_class_2': 'dst'})
  832. flow2['src'] = flow2['src'].map(lambda c: f"1: {c}")
  833. flow2['dst'] = flow2['dst'].map(lambda c: f"2: {c}")
  834. for _, r in flow2.iterrows():
  835. s = idx[r['src']]
  836. t = idx[r['dst']]
  837. source.append(s)
  838. target.append(t)
  839. value.append(r['count'])
  840. link_colors.append(node_colors[s].replace('rgb', 'rgba').replace(')', ',0.5)'))
  841. # flow3 → “3:Class”
  842. flow3 = flow3.rename(columns={'output_super_class_2': 'src', 'output_super_class_3': 'dst'})
  843. flow3['src'] = flow3['src'].map(lambda c: f"2: {c}")
  844. flow3['dst'] = flow3['dst'].map(lambda c: f"3: {c}")
  845. for _, r in flow3.iterrows():
  846. s = idx[r['src']]
  847. t = idx[r['dst']]
  848. source.append(s)
  849. target.append(t)
  850. value.append(r['count'])
  851. link_colors.append(node_colors[s].replace('rgb', 'rgba').replace(')', ',0.5)'))
  852. # 6) Compute hover info
  853. incoming = dict.fromkeys(nodes, 0)
  854. outgoing = dict.fromkeys(nodes, 0)
  855. for s_, t_, v_ in zip(source, target, value):
  856. outgoing[nodes[s_]] += v_
  857. incoming[nodes[t_]] += v_
  858. customdata = [f"Incoming: {incoming[n]}<br>Outgoing: {outgoing[n]}" for n in nodes]
  859. # 7) X positions
  860. x0 = [0.0] * len(col1)
  861. x1 = [0.33] * len(col2)
  862. x2 = [0.66] * len(col3)
  863. x3 = [1.0] * len(col4)
  864. x = x0 + x1 + x2 + x3
  865. # 8) Y positions
  866. y = []
  867. for col in (col1, col2, col3, col4):
  868. n = len(col)
  869. if n == 1:
  870. y.append(0.5)
  871. else:
  872. y.extend(list(np.linspace(0, 1, n)))
  873. # 9) **Return the Sankey trace directly** (no go.Trace wrapping)
  874. sankey = go.Sankey(
  875. name=title,
  876. arrangement='snap',
  877. node=dict(
  878. label=nodes,
  879. x=x,
  880. y=y,
  881. color=node_colors,
  882. pad=8,
  883. thickness=12,
  884. line=dict(color='black', width=0.3),
  885. customdata=customdata,
  886. hovertemplate='%{customdata}<extra>%{label}</extra>'
  887. ),
  888. link=dict(source=source, target=target, value=value, color=link_colors)
  889. )
  890. return sankey
  891. # %%
  892. # Now stack all 13 Sankey traces in a single portrait figure:
  893. if 'find_project_root' not in globals():
  894. from pathlib import Path
  895. def find_project_root() -> Path:
  896. current = Path.cwd().resolve()
  897. for candidate in (current, *current.parents):
  898. if (candidate / 'flywire_data').exists() and (candidate / 'input').exists():
  899. return candidate
  900. raise FileNotFoundError('Could not find the Giakoumas-et-al project root from the current working directory.')
  901. PROJECT_ROOT = find_project_root()
  902. INPUT_DIR = PROJECT_ROOT / 'input'
  903. OUTPUT_DIR = PROJECT_ROOT / 'output' / 'figures' / 'fig_S4'
  904. OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
  905. if 'connections' not in globals() or 'classification' not in globals():
  906. DATA_DIR = PROJECT_ROOT / 'flywire_data'
  907. CONNECTIONS_PATH = DATA_DIR / 'connections.csv'
  908. if not CONNECTIONS_PATH.exists():
  909. CONNECTIONS_PATH = DATA_DIR / 'connections.csv.zip'
  910. connections = pd.read_csv(CONNECTIONS_PATH)
  911. classification = pd.read_csv(DATA_DIR / 'classification.csv.gz')[['root_id', 'super_class']]
  912. # (A) Gather all dataframes:
  913. sets_phn_sa = {
  914. f'PhN-SA_v2_{i}': pd.read_csv(INPUT_DIR / 'PhN' / f'set_{i}.csv')
  915. for i in range(1, 7)
  916. }
  917. sets_pso = {
  918. 'DCSO': pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_1.csv'),
  919. 'aPhN1': pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_2.csv'),
  920. 'aPhN2': pd.read_csv(INPUT_DIR / 'PSO_SA' / 'set_3.csv'),
  921. }
  922. sets_mxlbn = {
  923. 'Sugar/Water': pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'sugar_water_GRNs.csv'),
  924. 'Bitter': pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'bitter_GRNs.csv'),
  925. 'Ir94e': pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'Ir94e_GRNs.csv'),
  926. 'Taste Peg': pd.read_csv(INPUT_DIR / 'MxLbN-SA' / 'taste_peg_GRNs.csv'),
  927. }
  928. all_sets = []
  929. all_sets += list(sets_phn_sa.items())
  930. all_sets += list(sets_pso.items())
  931. all_sets += list(sets_mxlbn.items())
  932. n_plots = len(all_sets) # 13
  933. # (B) Create subplots (13 rows × 1 column, each type="sankey")
  934. fig = make_subplots(
  935. rows = n_plots,
  936. cols = 1,
  937. shared_xaxes = False,
  938. shared_yaxes = False,
  939. vertical_spacing = 0.02,
  940. specs = [[{"type": "sankey"}] for _ in range(n_plots)],
  941. )
  942. # (C) Add each Sankey trace into its own row
  943. for i, (label, df) in enumerate(all_sets, start=1):
  944. sankey_trace = make_sankey_trace(
  945. grn_df = df,
  946. title = label,
  947. connections = connections,
  948. classification = classification,
  949. min_syn = 5
  950. )
  951. fig.add_trace(sankey_trace, row=i, col=1)
  952. # (D) Layout tweaks for portrait
  953. fig.update_layout(
  954. height = 300 * n_plots, # Still 3900 px for 13 plots
  955. width = 534, # One third of original 1600 px
  956. margin = dict(l=20, r=20, t=40, b=20),
  957. font = dict(size=10),
  958. title = "All Sankey Panels in One Portrait Figure",
  959. )
  960. fig.show()
  961. # To export as a single SVG:
  962. fig.write_image(OUTPUT_DIR / 'all_sankeys_portrait.svg')
  963. # %%

Full_hops_figure.ipynb at commit 9a82bcc, no license · at the source

Overview

Authors: Dimitrios S Giakoumas1, Julia M Zhu1, Alaina Jamal2, Zepeng Yao1,3,4,5
ORCID iDs: Zepeng Yao
  1. Department of Biology, University of Florida, Gainesville, FL 32611, USA
  2. Pine Crest School, Fort Lauderdale, FL 33334, USA
  3. Florida Chemical Senses Institute, University of Florida, Gainesville, FL 32611, USA
  4. McKnight Brain Institute, University of Florida, Gainesville, FL 32611, USA
  5. Genetics Institute, University of Florida, Gainesville, FL 32611, USA
Institutions: University of Florida (United States)
Journal: iScience, volume 29, issue 9, article 117359
Dates: received 8 December 2025; accepted 12 August 2026; published online 28 August 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1016/j.isci.2026.117359 · PMID 42701363 · PMCID PMC13545389 · OpenAlex W7204565829
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: drosophila (organism), systems (subfield)
Methods: Machine learning, Statistics
Keywords: Drosophila, connectome, feeding, gut, pharyngeal, gustatory
Topic: Neurobiology and Insect Physiology Research (Cellular and Molecular Neuroscience, Neuroscience), according to OpenAlex
Funding: National Institutes of Health; NIDDK NIH HHS (R01 DK139131); National Institute of Diabetes and Digestive and Kidney Diseases (R01DK139131); University of Florida
Citations: not cited yet (Europe PMC); 86 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.

google/neuroglancer

License: Apache-2.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: da443d25610b23c40c4a46ed9b92129e6203e68e, 22 September 2026
Languages: TypeScript (702), Python (133), C++ (23), JavaScript (20), Shell (11), C/C++ (6), Go (3), NEURON (1), Jupyter (1), Rust (1), C (1)
Size: 2,060 files, 902 scripts
Software Heritage: archived
Found in: the resources table
Holds: README, license file, environment (pyproject.toml, setup.py, uv.lock, docs/pyproject.toml, python/Dockerfile, src/mesh/draco/Dockerfile, src/sliceview/compresso/Dockerfile, src/sliceview/crackle/Dockerfile, src/sliceview/jxl/compile.Dockerfile, src/sliceview/jxl/optimize.Dockerfile, src/sliceview/png/Dockerfile), tests, continuous integration, documentation, 1 notebook
Not found: CITATION.cff
Tools: NumPy (67 files), Pillow (3 files), SciPy (3 files), pandas (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
904 files

ZepengYao-Lab/Giakoumas-et-al

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 9a82bcc1877e1285ac3a6db3c2417b8bb507d161, 25 September 2026
Languages: Python (31), Jupyter (19), R (5)
Size: 152 files, 55 scripts
Software Heritage: not archived
Found in: the resources table
Holds: README, CITATION.cff, environment (pyproject.toml, giakoumas-connectome-python/pyproject.toml), tests, 19 notebooks
Not found: license file, continuous integration, documentation
Tools: pandas (25 files), Matplotlib (14 files), NumPy (10 files), seaborn (6 files), SciPy (5 files), statsmodels (5 files), tidyverse (5 files), Plotly (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
56 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 957 scripts, each with its path and the digest of its content;
  • 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Code and data availability statement

The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it says that the data are available on request
  • it says that the code is available on request

Read it in the paper: doi.org/10.1016/j.isci.2026.117359.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Authors: added Zepeng Yao (0000-0002-0186-5503); removed Zepeng Yao

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 6 keywords, 4 funders, 85 references.

Cite

This paper

Giakoumas, D. S., Zhu, J. M., Jamal, A., & Yao, Z. (2026). Connectomic mapping of pharyngeal and gut sensory circuits in adult &lt;i&gt;Drosophila&lt;/i&gt;. iScience, 29(9), 117359. https://doi.org/10.1016/j.isci.2026.117359

BibTeX

@article{giakoumas2026connectomic,
author = {Giakoumas, Dimitrios S and Zhu, Julia M and Jamal, Alaina and Yao, Zepeng},
title = {{Connectomic mapping of pharyngeal and gut sensory circuits in adult \&lt;i\&gt;Drosophila\&lt;/i\&gt;}},
journal = {iScience},
year = {2026},
month = aug,
volume = {29},
number = {9},
pages = {117359},
publisher = {Elsevier},
issn = {2589-0042},
doi = {10.1016/j.isci.2026.117359},
url = {https://doi.org/10.1016/j.isci.2026.117359},
pmid = {42701363},
pmcid = {PMC13545389}
}

RIS

TY - JOUR
AU - Giakoumas, Dimitrios S
AU - Zhu, Julia M
AU - Jamal, Alaina
AU - Yao, Zepeng
TI - Connectomic mapping of pharyngeal and gut sensory circuits in adult &lt;i&gt;Drosophila&lt;/i&gt;
T2 - iScience
J2 - iScience
PY - 2026
DA - 2026/08/28
VL - 29
IS - 9
SP - 117359
SN - 2589-0042
PB - Elsevier
DO - 10.1016/j.isci.2026.117359
UR - https://doi.org/10.1016/j.isci.2026.117359
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.isci.2026.117359",
"type": "article-journal",
"title": "Connectomic mapping of pharyngeal and gut sensory circuits in adult &lt;i&gt;Drosophila&lt;/i&gt;",
"container-title": "iScience",
"author": [
{
"family": "Giakoumas",
"given": "Dimitrios S"
},
{
"family": "Zhu",
"given": "Julia M"
},
{
"family": "Jamal",
"given": "Alaina"
},
{
"family": "Yao",
"given": "Zepeng"
}
],
"container-title-short": "iScience",
"volume": "29",
"issue": "9",
"page": "117359",
"DOI": "10.1016/j.isci.2026.117359",
"PMID": "42701363",
"PMCID": "PMC13545389",
"ISSN": "2589-0042",
"publisher": "Elsevier",
"URL": "https://doi.org/10.1016/j.isci.2026.117359",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
28
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41586-026-10735-w [code]
Distributed control circuits across a brain-and-cord connectome.
Journal: Nature
In common: Plotly, seaborn, tidyverse, 4 other tools, drosophila, 23 references, author Zepeng Yao
[2] doi:10.1016/j.isci.2026.117467
Larval taste exposure creates compound-specific behavioral modifications in adult &lt;i&gt;Drosophila melanogaster&lt;/i&gt;.
Journal: iScience
In common: drosophila, 13 references
[3] doi:10.1016/j.isci.2026.116206 [code]
Gut distension evokes rapid neural dynamics in vagal and hindbrain populations of larval zebrafish.
Journal: iScience
In common: Pillow, statsmodels, seaborn, 4 other tools, systems, 7 references
[4] doi:10.3389/fnsys.2026.1822122 [code]
Convergence-divergence circuits for multimodal integration of innate and learned opponent valences.
Journal: Frontiers in systems neuroscience
In common: Plotly, Pillow, seaborn, 4 other tools, systems, 7 references
[5] doi:10.1038/s41467-026-72152-x [code]
Centralized brain networks controlling antennal grooming coordination.
Journal: Nature communications
In common: statsmodels, seaborn, pandas, 3 other tools, drosophila, systems, 7 references
[6] doi:10.1371/journal.pbio.3003959 [code]
Recurrent synapses between CO2-sensitive olfactory sensory neurons enable robust CO2 detection in Aedes aegypti mosquitoes.
Journal: PLoS biology
In common: statsmodels, seaborn, pandas, 3 other tools, drosophila, 7 references
[7] doi:10.1371/journal.pcbi.1014657 [code]
Population morphology implies a common developmental blueprint for Drosophila motion detectors.
Journal: PLoS computational biology
In common: statsmodels, seaborn, pandas, 3 other tools, drosophila, 4 references
[8] doi:10.1038/s41586-026-10797-w [code]
A global molecular code for birth order and neuronal identity in Drosophila.
Journal: Nature
In common: Pillow, tidyverse, pandas, 2 other tools, drosophila, 4 references
[9] doi:10.1038/s41593-026-02388-9 [code]
Hippocampal CA3 connectomics reveals a gradient of mossy fiber inputs and selective feedforward inhibition onto pyramidal cells.
Journal: Nature neuroscience
In common: Plotly, Pillow, seaborn, 4 other tools, systems, 2 references
[10] doi:10.1038/s41467-026-72437-1 [code]
High-speed whole-brain imaging in Drosophila.
Journal: Nature communications
In common: Plotly, Pillow, statsmodels, 4 other tools, drosophila, systems, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.