OSCR

An integrated single cell and spatial omics atlas of human prenatal development

Code ↔ Paper

19 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 19 matches
  1. [1] § Methods › Whole embryo: scRNAseq integration, batch correction and visualisation ↔ reproducibility/modules/module_generate/submodule_synthetic_simple/example_operations.ipynb, lines 292–373 · score 0.84 · n_neighbors, sc.pp.neighbors, sc.tl.umap, Scanpy, dimensionality, heatmaps
  2. [2] § Methods › Community data: Realignment and gene space harmonisation ↔ idtrack/_harmonize_features.py, lines 22–82 · score 0.81 · failed conversions, Gene symbols, gene identifiers, IDTrack, conflicts, audit
  3. [3] § Methods › Whole embryo: scRNAseq integration, batch correction and visualisation ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 1793–1810 · score 0.81 · n_neighbors, sc.pp.neighbors, sc.tl.umap, Scanpy, scVI, atlas
  4. [4] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › scTRAM: Trajectory-specific fidelity evaluation ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/incremental_analysis.ipynb, lines 70–155 · score 0.81 · graph edit distance, spectral distance, trustworthiness, Spearman, concordance, stress
  5. [5] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › Integrated benchmarking framework ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/helper/suo_scib_iterative.py, lines 25–133 · score 0.79 · min max scaled, scIB, aggregate score, batch correction, conservation, trajectory
  6. [6] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 198–252 · score 0.78 · TarDis, scPoli, scTRAM, scVI, Harmony, PCA
  7. [7] § Methods › Community data: Integration with scVI ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 198–252 · score 0.76 · TarDis, scPoli, scTRAM, scVI, Harmony, PCA
  8. [8] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › Integrated benchmarking framework ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 254–290 · score 0.75 · scTRAM, scIB, aggregate score, batch correction, topology, conservation
  9. [9] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › scIB: Conventional integration quality assessment ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/incremental_analysis.ipynb, lines 241–324 · score 0.70 · graph connectivity, scIB, cLISI, iLISI, isolated, silhouette
  10. [10] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › scTRAM: Trajectory-specific fidelity evaluation ↔ reproducibility/modules/module_evaluate/example_operations_adjacency.ipynb, lines 92–119 · score 0.68 · graph edit distance, spectral distance, adjacency, correlations, scores, metric
  11. [11] § Methods › HDCA: Integration and annotation harmonisation of the final integrated community and whole embryo data ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/helper/model_training.py, lines 15–97 · score 0.66 · gene likelihood, batch_key, trained, covariate, latent, layer
  12. [12] § Results › Benchmarking atlas integration with trajectory fidelity ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 380–453 · score 0.66 · TarDis, scPoli, scVI, Harmony, PCA, scTRAM
  13. [13] § Methods › scRNAseq gene expression and UMAP visualisation ↔ reproducibility/modules/module_generate/submodule_synthetic_simple/example_operations.ipynb, lines 292–373 · score 0.64 · sc.pl.umap, Scanpy, pp, heatmaps, matrices
  14. [14] § Methods › HDCA: Integration and annotation harmonisation of the final integrated community and whole embryo data ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/incremental_analysis.ipynb, lines 70–155 · score 0.63 · Normalised Mutual Information, macrophage, myeloid, concordance, Suo, neighbour
  15. [15] § Results › Benchmarking atlas integration with trajectory fidelity ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 1793–1810 · score 0.58 · ground truth, UMAP, HSC, MPP, scPoli, scTRAM
  16. [16] § Methods › Community data: Highly variable genes selection pipeline ↔ idtrack/_harmonize_features.py, lines 22–82 · score 0.53 · harmonised genes, IDTrack, single cell, logs, conversion, downstream
  17. [17] § Methods › Community data: Public human developmental single cell data acquisition and preprocessing ↔ idtrack/_harmonize_features.py, lines 876–941 · score 0.52 · trivial, age, provenance, donor, concatenated, schema
  18. [18] § Methods › Community data: Integration with scVI ↔ idtrack/_harmonize_features.py, lines 943–979 · score 0.51 · AnnData, Scanpy, uns, post, validation, genes
  19. [19] § Methods › Whole embryo: Xenium preprocessing and analysis ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/helper/model_training.py, lines 15–97 · score 0.50 · categorical covariate, unlabelled, SCANVI, latent, layers, Batch

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 1,813 lines · 57 KB · BSD-3-Clause · 6 matches

  1. # %% [markdown]
  2. # # Imports
  3. # %%
  4. 0
  5. # %%
  6. import sys
  7. import os
  8. import gc
  9. import networkx as nx
  10. import numpy as np
  11. import pandas as pd
  12. import tqdm
  13. # %%
  14. import anndata as ad
  15. import scanpy as sc
  16. # import scvelo as scv
  17. # import cellrank as cr
  18. # import scanpy.external as sce
  19. # scv.settings.verbosity = 3
  20. # cr.settings.verbosity = 2
  21. # sc.settings.verbose = 3
  22. # %%
  23. working_directory = "/Users/kemalinecik/git_nosync/sctram"
  24. sys.path.append(working_directory)
  25. sys.path.append("/Users/kemalinecik/git_nosync/sctram/reproducibility/server_sync/experiments")
  26. from incremental_training.helper.constants import metrics_scib, metric_direction_dict
  27. from development_atlas_figure.helper.get_hdca_trainings_with_paths import get_hdca_trainings_with_paths
  28. from development_atlas_figure.helper.pickle_operations import save_pickle_bundle, load_pickle_bundle
  29. from development_atlas_figure.helper.csv_operations import (
  30. harmonize_union_with_report,
  31. analyze_column_consistency,
  32. combine_dataframes_on_columns
  33. )
  34. from sctram.api._lower_level import TrajectoryEvaluationAPI
  35. from sctram.generate.real import sc_suo_developmental_complete
  36. from sctram.input import InputTrajectories
  37. # %%
  38. %matplotlib inline
  39. %config InlineBackend.figure_format='retina'
  40. import pickle
  41. import matplotlib.pyplot as plt
  42. import seaborn as sns
  43. import colorcet as cc
  44. import matplotlib.patheffects as path_effects
  45. from adjustText import adjust_text
  46. from matplotlib import gridspec
  47. from networkx.drawing.nx_agraph import graphviz_layout
  48. import matplotlib as mpl
  49. import matplotlib.patches as mpatches
  50. from matplotlib.ticker import FormatStrFormatter
  51. from adjustText import adjust_text
  52. import matplotlib.patheffects as path_effects
  53. import matplotlib.cm as cm
  54. from plottable import ColumnDefinition, Table
  55. from plottable.cmap import normed_cmap
  56. from matplotlib.ticker import MaxNLocator
  57. from plottable.plots import bar
  58. _rcparams_path = os.path.join(working_directory, "reproducibility/figure_rcparams/rcparams.pickle")
  59. with open(_rcparams_path, "rb") as file:
  60. _rcparams = pickle.load(file)
  61. plt.rcParams.update(_rcparams)
  62. plt.rcParams["font.family"] = "Arial"
  63. plt.rcParams["font.sans-serif"] = ["Arial"]
  64. # %%
  65. temp_move = "/Users/kemalinecik/Downloads/temp_move"
  66. # %% [markdown]
  67. # how many
  68. # %%
  69. (
  70. df_hdca_component_adata_paths,
  71. df_hdca_lvl3_adata_paths,
  72. out_df,
  73. ground_truth_trajectories,
  74. ground_truth_trajectories_annotation_level,
  75. trajectory_subset_dict
  76. ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
  77. ground_truth_trajectories.keys()
  78. # %%
  79. InputTrajectories(ground_truth_trajectories).get_trajectory("germline", False)
  80. # %%
  81. InputTrajectories(ground_truth_trajectories).get_trajectory("Haem_ref", False)
  82. # %%
  83. InputTrajectories(ground_truth_trajectories).get_trajectory("pns_neuro", False)
  84. # %%
  85. fig, ax = InputTrajectories(ground_truth_trajectories).get_trajectory("germline", False).plot_trajectory(
  86. title="",
  87. figsize=(12, 1.65),
  88. label_offset=0,
  89. return_ax=True
  90. )
  91. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/graph_germline.pdf")
  92. # %%
  93. fig, ax = InputTrajectories(ground_truth_trajectories).get_trajectory("Haem_ref", False).plot_trajectory(
  94. title="",
  95. figsize=(12, 7),
  96. label_offset=0,
  97. return_ax=True
  98. )
  99. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/graph_Haem_ref.pdf")
  100. # %%
  101. fig, ax = InputTrajectories(ground_truth_trajectories).get_trajectory("pns_neuro", False).plot_trajectory(
  102. title="",
  103. figsize=(12, 2),
  104. label_offset=0,
  105. return_ax=True
  106. )
  107. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/graph_pns_neuro.pdf")
  108. # %% [markdown]
  109. # # Figure
  110. # %%
  111. def logistic_invert_columns(df, col=None, k=1.0):
  112. """
  113. Apply logistic inversion scaling to a column or entire DataFrame.
  114. Parameters
  115. ----------
  116. df : pd.DataFrame or pd.Series
  117. Input DataFrame or Series.
  118. col : str or None, default=None
  119. Column name to transform. If None, apply to all numeric columns.
  120. k : float, default=1.0
  121. Steepness of the sigmoid. Larger k -> sharper curve.
  122. Returns
  123. -------
  124. pd.DataFrame or pd.Series
  125. Transformed DataFrame/Series with values in (0, 1).
  126. """
  127. if isinstance(df, pd.Series): # if a single column (Series)
  128. return 1 / (1 + np.exp(k * (df - df.mean())))
  129. if col is not None: # single column of DataFrame
  130. df = df.copy()
  131. x = df[col]
  132. df[col] = 1 / (1 + np.exp(k * (x - x.mean())))
  133. return df
  134. # all numeric columns
  135. df = df.copy()
  136. for c in df.select_dtypes(include=np.number).columns:
  137. x = df[c]
  138. df[c] = 1 / (1 + np.exp(k * (x - x.mean())))
  139. return df
  140. def logistic_invert_rowwise(df, k=1.0):
  141. """
  142. Apply logistic inversion scaling row-wise to all numeric columns.
  143. Parameters
  144. ----------
  145. df : pd.DataFrame
  146. Input DataFrame.
  147. k : float, default=1.0
  148. Steepness of the sigmoid. Larger k -> sharper curve.
  149. Returns
  150. -------
  151. pd.DataFrame
  152. Transformed DataFrame with values in (0, 1).
  153. """
  154. df = df.copy()
  155. numeric_cols = df.select_dtypes(include=np.number).columns
  156. # Apply rowwise
  157. def transform_row(row):
  158. mean_val = row[numeric_cols].mean()
  159. return 1 / (1 + np.exp(k * (row[numeric_cols] - mean_val)))
  160. df[numeric_cols] = df[numeric_cols].apply(transform_row, axis=1)
  161. return df
  162. # %%
  163. def rank_metrics(df, metric_direction_dict, metric_index_pos, method='min'):
  164. ranked_df = pd.DataFrame(index=df.index, columns=df.columns, dtype=float)
  165. for indices, row in df.iterrows():
  166. metric = indices[metric_index_pos]
  167. direction = metric_direction_dict.get(metric)
  168. if direction not in ('increasing', 'decreasing'):
  169. raise ValueError
  170. numeric_row = pd.to_numeric(row, errors='coerce')
  171. if numeric_row.isna().any():
  172. ranked_df.loc[indices] = np.nan
  173. continue
  174. if direction == 'increasing':
  175. ranks = numeric_row.rank(method='dense', ascending=False)
  176. else:
  177. ranks = numeric_row.rank(method='dense', ascending=True)
  178. # ranks = 1 - pd.Series(MinMaxScaler().fit_transform(row.values.reshape(-1, 1)).flatten(), index=row.index)
  179. ranked_df.loc[indices] = ranks.astype('Float64')
  180. return ranked_df
  181. result_df = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_result_df_dataframe.pkl"))
  182. df_pivot = result_df.pivot(index=["trajectory_class", 'trajectory', 'path', 'metric'], columns='representation', values='score')
  183. df_pivot_rank = rank_metrics(df_pivot, metric_direction_dict, metric_index_pos=3, method='min')
  184. df_pivot_rank_mean = df_pivot_rank.groupby(level=1).mean()
  185. rename_model_name_dict = {
  186. 'pca': 'PCA',
  187. 'X_pca': 'PCA',
  188. 'harmony': 'Harmony',
  189. 'invae': 'inVAE',
  190. 'scanvi': 'scANVI',
  191. 'scvi': 'scVI',
  192. 'scpoli': 'scPoli',
  193. 'scanoroma': 'Scanoroma',
  194. 'tardis_1': 'DELETE ME',
  195. 'tardis_2': 'TarDis',
  196. 'tardis': 'TarDis',
  197. }
  198. col_order = ['Scanoroma', 'Harmony', 'TarDis', 'PCA', 'scPoli', 'scVI']
  199. sctram_summary = (logistic_invert_rowwise(df_pivot_rank).groupby(level=[2, 3]).mean().groupby(level=0).mean())
  200. _sctram_summary = (pd.DataFrame(logistic_invert_rowwise(df_pivot_rank).groupby(level=3).mean().mean(), columns=["Total scTRAM"]).T)
  201. sctram_summary = pd.concat([sctram_summary, _sctram_summary], axis=0)
  202. sctram_summary = sctram_summary.rename(
  203. index={"adjacency": "Adjacency-based", "embedding": "Embedding-based", "pseudotime": "Pseudotime-based"},
  204. columns=rename_model_name_dict,
  205. )[col_order]
  206. sctram_summary.index.name = None
  207. sctram_summary.columns.name = "Representation"
  208. sctram_summary
  209. # %%
  210. # sctram_summary = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_sctram_summary_dataframe.pkl")) # invert back to rank again
  211. scib_summary = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_scib_summary_dataframe.pkl"))
  212. replacements = {
  213. "Pseudotime-based": "Cell Order\ncontinuity",
  214. "Embedding-based": "Trajectory\ndominance",
  215. "Adjacency-based": "Topology\nconservation",
  216. "Bio conservation": "Biological\nconservation",
  217. "Batch correction": "Batch\ncorrection",
  218. }
  219. summary = pd.concat([scib_summary, sctram_summary], keys=["scIB", "scTRAM"], axis=0)
  220. summary.index.names = ["Source", "Metric Group"]
  221. summary.columns.name = "Representation"
  222. new_index_tuples = [
  223. ("Aggregate score", metric.split("Total ")[1]) if metric.startswith("Total ") else (src, metric) for src, metric in summary.index
  224. ]
  225. summary.index = pd.MultiIndex.from_tuples(new_index_tuples, names=summary.index.names)
  226. weights = { # just provisional for now
  227. "scTRAM": 0.25,
  228. "scIB": 0.75,
  229. }
  230. row_ts = summary.loc[("Aggregate score", "scTRAM")]
  231. row_si = summary.loc[("Aggregate score", "scIB")]
  232. combined = weights["scTRAM"] * row_ts + weights["scIB"] * row_si
  233. combined_df = pd.DataFrame(combined).T
  234. combined_df.index = pd.MultiIndex.from_tuples([("Aggregate score", "Total")], names=summary.index.names)
  235. summary = pd.concat([summary, combined_df])
  236. summary = summary.sort_index(ascending=False)
  237. summary = summary.sort_values(by=[("Aggregate score", "Total")], axis=1, ascending=False).astype(np.float64)
  238. summary.rename(index=replacements, level="Metric Group", inplace=True)
  239. summary_T = summary.T
  240. summary
  241. # %%
  242. df_plot = summary_T.reset_index().copy()
  243. df_plot_dict = {f"{lvl0}\n{lvl1}": (lvl0, lvl1) for (lvl0, lvl1) in df_plot.columns}
  244. df_plot.columns = [f"{lvl0}\n{lvl1}" for (lvl0, lvl1) in df_plot.columns]
  245. cmap_fn = lambda col_data: normed_cmap(col_data, cmap=cm.PRGn, num_stds=2.5)
  246. col_defs = []
  247. for col in df_plot.columns:
  248. # Create a “normalized color map” for this column:
  249. lvl0, lvl1 = df_plot_dict[col]
  250. if lvl0 == "Aggregate score" or lvl0 == "Representation":
  251. continue
  252. cmap = cmap_fn(df_plot[col])
  253. cd = ColumnDefinition(
  254. name=col, # Must exactly match the flattened column name in df_plot
  255. width=0.4, # width of each column in inches (tweak as needed)
  256. cmap=cmap, # cell background colormap (normed by this column’s values)
  257. text_cmap=cmap, # text color also follows same colormap
  258. textprops={
  259. "ha": "center",
  260. "bbox": {"boxstyle": "circle", "pad": 0},
  261. },
  262. group=lvl0,
  263. title=lvl1,
  264. formatter="{:.3f}".format, # three decimal places
  265. )
  266. col_defs.append(cd)
  267. i = 0
  268. for col in df_plot.columns:
  269. # Create a “normalized color map” for this column:
  270. lvl0, lvl1 = df_plot_dict[col]
  271. if lvl0 != "Aggregate score" or lvl0 == "Representation":
  272. continue
  273. cmap = cmap_fn(df_plot[col])
  274. i += 1
  275. c_min = min(df_plot[col])
  276. c_max = max(df_plot[col])
  277. cmap = cmap_fn(df_plot[col])
  278. cd = ColumnDefinition(
  279. name=col,
  280. width=0.3,
  281. plot_fn=bar,
  282. plot_kw={
  283. "cmap": cmap, #cm.YlGnBu,
  284. "plot_bg_bar": False,
  285. # "xlim": [c_min, c_max],
  286. "annotate": True,
  287. "height": 0.7,
  288. "formatter": "{:.3f}",
  289. "textprops": {"fontsize": 10}
  290. },
  291. group=lvl0,
  292. title=lvl1,
  293. border="left" if i == 0 else None,
  294. )
  295. col_defs.append(cd)
  296. col_defs += [ColumnDefinition(name="Representation\n", title="", width=0.3, textprops={"ha": "left", "weight": "bold"})]
  297. # Allow to manipulate text post-hoc (in illustrator)
  298. with mpl.rc_context({"svg.fonttype": "none"}):
  299. fig, ax = plt.subplots(figsize=(len(summary_T.columns) * 1.3, 3 + 0.25 * len(summary_T.index)))
  300. tab = Table(
  301. df_plot,
  302. cell_kw={
  303. "linewidth": 0,
  304. "edgecolor": "k",
  305. },
  306. row_dividers=True,
  307. footer_divider=True,
  308. textprops={"fontsize": 10, "ha": "center"},
  309. row_divider_kw={"linewidth": 0.25, "linestyle": "-"},
  310. col_label_divider_kw={"linewidth": 1, "linestyle": "-"},
  311. column_border_kw={"linewidth": 1, "linestyle": "-"},
  312. column_definitions=col_defs,
  313. even_row_color="#97ffff",
  314. odd_row_color="#ffffff",
  315. ax=ax,
  316. index_col="Representation\n",
  317. ).autoset_fontcolors(colnames=df_plot.columns)
  318. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/table.pdf")
  319. # %% [markdown]
  320. # # Figure
  321. # %%
  322. result_df = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_result_df_dataframe.pkl"))
  323. df_pivot = result_df.pivot(index=["trajectory_class", 'trajectory', 'path', 'metric'], columns='representation', values='score')
  324. df_pivot_rank = rank_metrics(df_pivot, metric_direction_dict, metric_index_pos=3, method='min')
  325. df_pivot_rank_mean = logistic_invert_rowwise(df_pivot_rank).groupby(level=1).mean()
  326. rename_model_name_dict = {
  327. 'pca': 'PCA',
  328. 'X_pca': 'PCA',
  329. 'harmony': 'Harmony',
  330. 'invae': 'inVAE',
  331. 'scanvi': 'scANVI',
  332. 'scvi': 'scVI',
  333. 'scpoli': 'scPoli',
  334. 'scanoroma': 'Scanoroma',
  335. 'tardis_1': 'DELETE ME',
  336. 'tardis_2': 'TarDis',
  337. 'tardis': 'TarDis',
  338. }
  339. col_order = ['Scanoroma', 'Harmony', 'TarDis', 'PCA', 'scPoli', 'scVI']
  340. def _(df):
  341. df = df.rename(columns=rename_model_name_dict)
  342. df = df[col_order]
  343. fig_height=4
  344. n_cols = df.shape[1]
  345. col_width = 0.7 # Width per column in inches
  346. fig_width = n_cols * col_width
  347. # fig_height = 8 # Fixed height for better vertical readability
  348. # Create figure with dynamic width
  349. plt.figure(figsize=(fig_width, fig_height))
  350. # Create boxplot with refined styling
  351. bp = sns.boxplot(data=df,
  352. width=0.8,
  353. showfliers=False,
  354. boxprops=dict(facecolor='lightgray', edgecolor='black', linewidth=0.81),
  355. whiskerprops=dict(color='black', linewidth=0.81),
  356. medianprops=dict(color='darkred', linewidth=0.81),
  357. capprops=dict(color='black', linewidth=0.81),
  358. saturation=1,
  359. zorder=1)
  360. # Add swarmplot with sophisticated aesthetics
  361. sp = sns.swarmplot(data=df,
  362. size=4,
  363. palette='dark:black',
  364. edgecolor='white',
  365. linewidth=0.6,
  366. alpha=0.8,
  367. marker='o',
  368. zorder=2)
  369. # Enhance axis styling
  370. plt.xticks(rotation=90, ha='center', fontsize=10)
  371. plt.yticks(fontsize=10)
  372. plt.xlabel('Representation', fontsize=10, labelpad=10)
  373. plt.ylabel("Performance", fontsize=10, labelpad=10)
  374. # plt.gca().set_ylim(1.75, 5.5)
  375. plt.gca().yaxis.set_major_formatter(FormatStrFormatter('%.2f'))
  376. # Add light grid and adjust borders
  377. bp.grid(axis='y', linestyle='--', alpha=0.4)
  378. sns.despine(left=True, trim=False)
  379. # Adjust layout with professional spacing
  380. plt.tight_layout(rect=[0.05, 0.05, 0.95, 0.95])
  381. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/modelcomp_boxplot.pdf")
  382. _(df_pivot_rank_mean)
  383. # %%
  384. df_pivot_rank_mean = logistic_invert_rowwise(df_pivot_rank).groupby(level=[0, 1]).mean()
  385. df_pivot_rank_mean = df_pivot_rank_mean.rename(columns=rename_model_name_dict)
  386. df_pivot_rank_mean = df_pivot_rank_mean[col_order]
  387. df_long = df_pivot_rank_mean.reset_index().melt(id_vars=['trajectory_class', 'trajectory'], var_name='representation', value_name='score')
  388. # %%
  389. # match boxplot’s height & per-column width
  390. fig_height = 3
  391. n_cols = df_long["representation"].nunique()
  392. col_width = 1.4
  393. fig_width = n_cols * col_width
  394. n_tc = len(np.unique(df_long["trajectory_class"]))
  395. fig, ax = plt.subplots(figsize=(fig_width, fig_height))
  396. # draw bars
  397. bp = sns.barplot(
  398. data=df_long,
  399. x="representation",
  400. y="score",
  401. hue="trajectory_class",
  402. palette=sns.color_palette("Greys", n_colors=n_tc+2),
  403. width=0.8,
  404. saturation=1,
  405. errorbar="sd",
  406. capsize=0.05,
  407. err_kws={'color': 'black', 'linewidth': 0.81},
  408. edgecolor="black",
  409. linewidth=0.81,
  410. zorder=1,
  411. ax=ax
  412. )
  413. # enforce edge styling
  414. for patch in bp.patches:
  415. patch.set_edgecolor("black")
  416. patch.set_linewidth(0.81)
  417. # axes
  418. ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
  419. ax.tick_params(axis="y", labelsize=10)
  420. ax.set_xlabel('Representation', fontsize=10, labelpad=10)
  421. plt.ylabel("Performance", fontsize=10, labelpad=10)
  422. ax.set_ylim(0.1, None)
  423. ax.yaxis.set_major_formatter(FormatStrFormatter('%.2f'))
  424. # grid + despine
  425. ax.grid(axis='y', linestyle='--', alpha=0.4)
  426. sns.despine(left=True, trim=False)
  427. # remove the old legend
  428. ax.legend_.remove()
  429. # build new, professional legend
  430. handles, labels = ax.get_legend_handles_labels()
  431. slim_handles = [
  432. mpatches.Patch(
  433. facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
  434. edgecolor="black",
  435. linewidth=0.5 # thinner border
  436. )
  437. for h in handles
  438. ]
  439. leg = ax.legend(
  440. slim_handles,
  441. labels,
  442. title="Trajectory Group",
  443. loc="center left",
  444. bbox_to_anchor=(1.02, 0.5),
  445. frameon=False,
  446. alignment="left",
  447. title_fontsize=8,
  448. fontsize=8,
  449. borderaxespad=0.5,
  450. handletextpad=0.4,
  451. labelspacing=0.5,
  452. )
  453. # tighten layout (accounts for legend)
  454. plt.tight_layout(rect=[0, 0, 0.85, 1])
  455. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/modelcomp_barplot1.pdf")
  456. # %%
  457. df_pivot_rank_mean = logistic_invert_rowwise(df_pivot_rank).groupby(level=[2, 1]).mean()
  458. df_pivot_rank_mean = df_pivot_rank_mean.rename(columns=rename_model_name_dict)
  459. df_pivot_rank_mean = df_pivot_rank_mean[col_order]
  460. df_long = df_pivot_rank_mean.reset_index().melt(id_vars=['path', 'trajectory'], var_name='representation', value_name='score')
  461. # %%
  462. label_mapping = {
  463. "adjacency": 'Topology conservation',
  464. "embedding": 'Trajectory dominance',
  465. "pseudotime": 'Cell Order continuity',
  466. }
  467. # match boxplot’s height & per-column width
  468. fig_height = 3
  469. n_cols = df_long["representation"].nunique()
  470. col_width = 1.4
  471. fig_width = n_cols * col_width
  472. fig, ax = plt.subplots(figsize=(fig_width, fig_height))
  473. # draw bars
  474. bp = sns.barplot(
  475. data=df_long,
  476. x="representation",
  477. y="score",
  478. hue="path",
  479. palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
  480. width=0.8,
  481. saturation=1,
  482. errorbar="sd",
  483. capsize=0.05,
  484. err_kws={'color': 'black', 'linewidth': 0.81},
  485. edgecolor="black",
  486. linewidth=0.81,
  487. zorder=1,
  488. ax=ax
  489. )
  490. # enforce edge styling
  491. for patch in bp.patches:
  492. patch.set_edgecolor("black")
  493. patch.set_linewidth(0.81)
  494. # axes
  495. ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
  496. ax.tick_params(axis="y", labelsize=10)
  497. ax.set_xlabel('Representation', fontsize=10, labelpad=10)
  498. plt.ylabel("Performance", fontsize=10, labelpad=10)
  499. ax.set_ylim(0.1, None)
  500. ax.yaxis.set_major_formatter(FormatStrFormatter('%.2f'))
  501. # grid + despine
  502. ax.grid(axis='y', linestyle='--', alpha=0.4)
  503. sns.despine(left=True, trim=False)
  504. # remove the old legend
  505. ax.legend_.remove()
  506. # build new, professional legend
  507. handles, labels = ax.get_legend_handles_labels()
  508. new_labels = [label_mapping.get(l, l) for l in labels]
  509. slim_handles = [
  510. mpatches.Patch(
  511. facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
  512. edgecolor="black",
  513. linewidth=0.5 # thinner border
  514. )
  515. for h in handles
  516. ]
  517. leg = ax.legend(
  518. slim_handles,
  519. new_labels,
  520. title="Metric Group",
  521. loc="center left",
  522. bbox_to_anchor=(1.02, 0.5),
  523. alignment="left",
  524. frameon=False,
  525. title_fontsize=8,
  526. fontsize=8,
  527. borderaxespad=0.5,
  528. handletextpad=0.4,
  529. labelspacing=0.5,
  530. )
  531. # tighten layout (accounts for legend)
  532. plt.tight_layout(rect=[0, 0, 0.85, 1])
  533. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/modelcomp_barplot2.pdf")
  534. # %% [markdown]
  535. # # Figure 2
  536. # %%
  537. result_df_main = pd.read_pickle(os.path.join(temp_move, f"df_boots_fig2_calculations.pkl"))
  538. (
  539. df_hdca_component_adata_paths,
  540. df_hdca_lvl3_adata_paths,
  541. out_df,
  542. ground_truth_trajectories,
  543. ground_truth_trajectories_annotation_level,
  544. trajectory_subset_dict
  545. ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
  546. # %%
  547. from math import sqrt
  548. def rank_metrics(df, metric_direction_dict, metric_index_pos, method='min'):
  549. ranked_df = pd.DataFrame(index=df.index, columns=df.columns, dtype=float)
  550. for indices, row in df.iterrows():
  551. metric = indices[metric_index_pos]
  552. direction = metric_direction_dict.get(metric)
  553. if direction not in ('increasing', 'decreasing'):
  554. raise ValueError
  555. numeric_row = pd.to_numeric(row, errors='coerce')
  556. if numeric_row.isna().any():
  557. ranked_df.loc[indices] = np.nan
  558. continue
  559. if direction == 'increasing':
  560. ranks = numeric_row.rank(method='dense', ascending=False)
  561. else:
  562. ranks = numeric_row.rank(method='dense', ascending=True)
  563. # ranks = 1 - pd.Series(MinMaxScaler().fit_transform(row.values.reshape(-1, 1)).flatten(), index=row.index)
  564. ranked_df.loc[indices] = ranks.astype('Float64')
  565. return ranked_df
  566. def plot_trajectory(
  567. G,
  568. title="trajectory_name",
  569. figsize=None,
  570. font_size=10,
  571. node_size=300,
  572. node_color="#8DA0CB", # Modern muted blue
  573. cmap="viridis",
  574. color_nodes_by_position=False,
  575. # Edge styling parameters
  576. edge_width=1,
  577. edge_color="gray",
  578. color_edges_by_weight=False,
  579. edge_weight_attribute="weight",
  580. edge_cmap="plasma",
  581. edge_vmin=None,
  582. edge_vmax=None,
  583. edge_colorbar=True,
  584. edge_labels=False,
  585. edge_label_format=".2f",
  586. edge_label_font_size=8,
  587. edge_label_color="black",
  588. edge_label_offset=None,
  589. # Layout parameters
  590. label_offset=None,
  591. title_y=1.02,
  592. layout_args="-Grankdir=LR -Gnodesep=0.25 -Granksep=0.25",
  593. file_name="blabla"
  594. ):
  595. pos = nx.nx_agraph.graphviz_layout(G, prog="dot", args=layout_args)
  596. x_coords = [pos[node][0] for node in G.nodes()] if G.nodes() else []
  597. y_coords = [pos[node][1] for node in G.nodes()] if G.nodes() else []
  598. # Create actual figure with determined size
  599. fig = plt.figure(figsize=figsize)
  600. ax = fig.add_subplot(111)
  601. # Node coloring logic
  602. if color_nodes_by_position:
  603. try:
  604. min_x, max_x = min(x_coords), max(x_coords)
  605. color_vals = [(x - min_x) / (max_x - min_x) if max_x != min_x else 0.5 for x in x_coords]
  606. node_color = plt.get_cmap(cmap)(color_vals)
  607. except Exception as e:
  608. print(f"Warning: Positional node coloring failed - {e}. Using base color.")
  609. # Edge coloring logic
  610. if color_edges_by_weight:
  611. weights = [G[u][v].get(edge_weight_attribute, 1.0) for u, v in G.edges]
  612. # Separate NaN and valid weights
  613. valid_weights = [w for w in weights if not np.isnan(w)]
  614. if valid_weights:
  615. edge_vmin = edge_vmin if edge_vmin is not None else min(valid_weights)
  616. edge_vmax = edge_vmax if edge_vmax is not None else max(valid_weights)
  617. norm = mpl.colors.Normalize(vmin=edge_vmin, vmax=edge_vmax)
  618. cmap_edge = plt.get_cmap(edge_cmap)
  619. # Create edge colors, using gray with alpha=0.712 for NaN weights
  620. edge_colors = []
  621. edge_alphas = []
  622. for w in weights:
  623. if np.isnan(w):
  624. edge_colors.append("lightgray")
  625. edge_alphas.append(0.712)
  626. else:
  627. edge_colors.append(cmap_edge(norm(w)))
  628. edge_alphas.append(1.0)
  629. sm = plt.cm.ScalarMappable(norm=norm, cmap=cmap_edge)
  630. sm.set_array(valid_weights)
  631. else:
  632. edge_colors = "lightgray"
  633. edge_alphas = 0.712
  634. sm = None
  635. else:
  636. edge_colors = edge_color
  637. edge_alphas = 1.0
  638. sm = None
  639. # Label positioning offsets
  640. if label_offset is None:
  641. label_offset = 0
  642. if edge_label_offset is None:
  643. edge_label_offset = 0
  644. # Calculate final label positions
  645. label_pos = {n: (x, y + label_offset) for n, (x, y) in pos.items()}
  646. # Draw nodes and edges
  647. nx.draw_networkx_nodes(
  648. G, pos, ax=ax, node_size=node_size, node_color=node_color, edgecolors="black", linewidths=0.8, alpha=1.0
  649. )
  650. node_radius = sqrt(node_size) / 2.0
  651. if color_edges_by_weight and isinstance(edge_alphas, list):
  652. # Draw edges individually when we have different alpha values
  653. for i, (u, v) in enumerate(G.edges()):
  654. nx.draw_networkx_edges(
  655. G,
  656. pos,
  657. edgelist=[(u, v)],
  658. ax=ax,
  659. edge_color=[edge_colors[i]],
  660. width=edge_width,
  661. alpha=edge_alphas[i],
  662. connectionstyle="arc3",
  663. arrows=False,
  664. )
  665. else:
  666. nx.draw_networkx_edges(
  667. G,
  668. pos,
  669. ax=ax,
  670. edge_color=edge_colors,
  671. width=edge_width,
  672. alpha=edge_alphas if not isinstance(edge_alphas, list) else 1.0,
  673. connectionstyle="arc3",
  674. arrows=False,
  675. )
  676. # Edge label placement
  677. if edge_labels:
  678. for u, v in G.edges():
  679. if edge_weight_attribute not in G[u][v]:
  680. continue
  681. weight = G[u][v][edge_weight_attribute]
  682. x1, y1 = pos[u]
  683. x2, y2 = pos[v]
  684. dx, dy = x2 - x1, y2 - y1
  685. length = (dx**2 + dy**2) ** 0.5
  686. if length == 0:
  687. continue # Skip self-loops
  688. # Calculate perpendicular offset position
  689. mid_x = (x1 + x2) / 2 + (dy / length) * edge_label_offset
  690. mid_y = (y1 + y2) / 2 - (dx / length) * edge_label_offset
  691. ax.text(
  692. mid_x,
  693. mid_y,
  694. f"{weight:{edge_label_format}}",
  695. fontsize=edge_label_font_size,
  696. color=edge_label_color,
  697. ha="center",
  698. va="center",
  699. # bbox=dict(boxstyle="round", facecolor="white", alpha=0.8, edgecolor="none"),
  700. )
  701. # Edge weight colorbar
  702. if color_edges_by_weight and edge_colorbar and sm:
  703. cbar = plt.colorbar(sm, ax=ax, orientation="horizontal", shrink=0.3, aspect=40, pad=0.05)
  704. cbar.set_label("Edge Weight", fontsize=font_size)
  705. cbar.ax.tick_params(labelsize=font_size - 2)
  706. fig.suptitle(title, y=title_y, fontsize=font_size + 2, fontweight="bold", color="#333333")
  707. # Final layout adjustments
  708. plt.subplots_adjust(
  709. left=0.03, right=0.97, top=0.97, bottom=0.15 if (color_edges_by_weight and edge_colorbar) else 0.03
  710. )
  711. ax.set_axis_off()
  712. plt.margins(y=1) # Add 20% margin on all sides
  713. plt.savefig(f"/Users/kemalinecik/Downloads/fig2fig/{file_name}.pdf")
  714. # %%
  715. for trajectory_focus, result_df in result_df_main.groupby('trajectory'):
  716. if trajectory_focus == "Haem_ref":
  717. decompose_method = "2"
  718. input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{trajectory_focus}_litc_{decompose_method}.pkl")
  719. with open(input_trajectories_path_litc, 'rb') as _file:
  720. lit = pickle.load(_file)
  721. break
  722. df_edges = []
  723. for ind, row in result_df.iterrows():
  724. trj = lit.get_trajectory(row["booted_trajectory"], False)
  725. for from_cell, to_cell in trj.edges():
  726. row_copy = row.copy()
  727. row_copy["from_cell"] = from_cell
  728. row_copy["to_cell"] = to_cell
  729. df_edges.append(row_copy)
  730. df_edges = pd.DataFrame(df_edges)
  731. df_edges.reset_index(drop=True, inplace=True)
  732. # %%
  733. df_edges_pivot = df_edges.pivot(index=['path', 'booted_trajectory', 'trajectory', 'metric', 'from_cell', 'to_cell'], columns='representation', values='score')
  734. df_edges_pivot_rank = rank_metrics(df_edges_pivot, metric_direction_dict, metric_index_pos=3, method='min')
  735. df_edges_pivot_rank_mean = logistic_invert_rowwise(df_edges_pivot_rank).groupby(level=[4, 5]).mean()
  736. # %%
  737. from scipy.stats import mannwhitneyu, binomtest, wilcoxon
  738. def compare_scores_and_get_means(group):
  739. scpoli_scores = group['scpoli'].dropna()
  740. scvi_scores = group['scvi'].dropna()
  741. if len(scpoli_scores) < 2 or len(scvi_scores) < 2:
  742. raise ValueError
  743. # since your “scores” are actually ranks (ordinal) for the same boot replicate per edge, you have paired, ordinal data.
  744. # Perform the Mann-Whitney U test
  745. # stat, p_value = mannwhitneyu(scpoli_scores, scvi_scores, alternative='two-sided')
  746. # Wilcoxon signed-rank — more power if the distribution of paired differences is roughly symmetric; handle ties with zero_method="pratt"
  747. method = 'exact' if len(scpoli_scores) <= 25 else 'approx'
  748. p_value = wilcoxon(scpoli_scores, scvi_scores, zero_method='pratt', alternative='two-sided', method=method).pvalue
  749. # Statistically different: return means.
  750. return pd.Series({
  751. 'mean_scpoli': np.mean(scpoli_scores),
  752. 'mean_scvi': np.mean(scvi_scores),
  753. 'p_value': p_value
  754. })
  755. df_edges_pivot_rank_p = logistic_invert_rowwise(df_edges_pivot_rank).groupby(['from_cell', 'to_cell']).apply(compare_scores_and_get_means).reset_index()
  756. alpha = 0.05 / len(df_edges_pivot_rank_p)
  757. df_edges_pivot_rank_p['mean_scvi_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scvi'], np.nan)
  758. df_edges_pivot_rank_p['mean_scpoli_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scpoli'], np.nan)
  759. df_edges_pivot_rank_p
  760. # %%
  761. trj = InputTrajectories(ground_truth_trajectories).get_trajectory(trajectory_focus, False)
  762. representations = ["scvi", "scpoli"]
  763. for ind, row in df_edges_pivot_rank_p.iterrows():
  764. for repr in representations:
  765. trj[row["from_cell"]][row["to_cell"]][f"weight_{repr}"] = row[f"mean_{repr}_final"]
  766. # %%
  767. for i in ['weight_scpoli', 'weight_scvi']:
  768. plot_trajectory(
  769. trj,
  770. figsize=(4, 11),
  771. edge_weight_attribute=i,
  772. title="",
  773. edge_vmin=0.4,
  774. edge_vmax=0.7,
  775. color_edges_by_weight=True,
  776. node_size=150,
  777. node_color="lightgray",
  778. edge_color='black',
  779. edge_width=5,
  780. edge_labels=False,
  781. edge_label_font_size=5,
  782. edge_cmap='viridis',
  783. edge_colorbar=False,
  784. file_name=f"{trajectory_focus}_{i}"
  785. )
  786. # %%
  787. for trajectory_focus, result_df in result_df_main.groupby('trajectory'):
  788. if trajectory_focus == "germline":
  789. decompose_method = "2"
  790. input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{trajectory_focus}_litc_{decompose_method}.pkl")
  791. with open(input_trajectories_path_litc, 'rb') as _file:
  792. lit = pickle.load(_file)
  793. break
  794. df_edges = []
  795. for ind, row in result_df.iterrows():
  796. trj = lit.get_trajectory(row["booted_trajectory"], False)
  797. for from_cell, to_cell in trj.edges():
  798. row_copy = row.copy()
  799. row_copy["from_cell"] = from_cell
  800. row_copy["to_cell"] = to_cell
  801. df_edges.append(row_copy)
  802. df_edges = pd.DataFrame(df_edges)
  803. df_edges.reset_index(drop=True, inplace=True)
  804. df_edges_pivot = df_edges.pivot(index=['path', 'booted_trajectory', 'trajectory', 'metric', 'from_cell', 'to_cell'], columns='representation', values='score')
  805. df_edges_pivot_rank = rank_metrics(df_edges_pivot, metric_direction_dict, metric_index_pos=3, method='min')
  806. df_edges_pivot_rank_mean = logistic_invert_rowwise(df_edges_pivot_rank).groupby(level=[4, 5]).mean()
  807. df_edges_pivot_rank_p = logistic_invert_rowwise(df_edges_pivot_rank).groupby(['from_cell', 'to_cell']).apply(compare_scores_and_get_means).reset_index()
  808. alpha = 0.05 / len(df_edges_pivot_rank_p)
  809. df_edges_pivot_rank_p['mean_scvi_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scvi'], np.nan)
  810. df_edges_pivot_rank_p['mean_scpoli_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scpoli'], np.nan)
  811. trj = InputTrajectories(ground_truth_trajectories).get_trajectory(trajectory_focus, False)
  812. representations = ["scvi", "scpoli"]
  813. for ind, row in df_edges_pivot_rank_p.iterrows():
  814. for repr in representations:
  815. trj[row["from_cell"]][row["to_cell"]][f"weight_{repr}"] = row[f"mean_{repr}_final"]
  816. # %%
  817. for i in ['weight_scpoli', 'weight_scvi']:
  818. plot_trajectory(
  819. trj,
  820. figsize=(4, 0.75),
  821. edge_weight_attribute=i,
  822. title="",
  823. edge_vmin=0.4,
  824. edge_vmax=0.7,
  825. color_edges_by_weight=True,
  826. node_size=150,
  827. node_color="lightgray",
  828. edge_color='black',
  829. edge_width=5,
  830. edge_labels=False,
  831. edge_label_font_size=5,
  832. edge_cmap='viridis',
  833. edge_colorbar=False,
  834. file_name=f"{trajectory_focus}_{i}"
  835. )
  836. # %%
  837. def compare_scores_and_get_means(group):
  838. tardis_scores = group['tardis'].dropna()
  839. scvi_scores = group['scvi'].dropna()
  840. if len(tardis_scores) < 2 or len(scvi_scores) < 2:
  841. raise ValueError
  842. # since your “scores” are actually ranks (ordinal) for the same boot replicate per edge, you have paired, ordinal data.
  843. # Perform the Mann-Whitney U test
  844. # stat, p_value = mannwhitneyu(tardis_scores, scvi_scores, alternative='two-sided')
  845. # Wilcoxon signed-rank — more power if the distribution of paired differences is roughly symmetric; handle ties with zero_method="pratt"
  846. method = 'exact' if len(tardis_scores) <= 25 else 'approx'
  847. p_value = wilcoxon(tardis_scores, scvi_scores, zero_method='pratt', alternative='two-sided', method=method).pvalue
  848. # Statistically different: return means.
  849. return pd.Series({
  850. 'mean_tardis': np.mean(tardis_scores),
  851. 'mean_scvi': np.mean(scvi_scores),
  852. 'p_value': p_value
  853. })
  854. # %%
  855. for trajectory_focus, result_df in result_df_main.groupby('trajectory'):
  856. if trajectory_focus == "pns_neuro":
  857. decompose_method = "2"
  858. input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{trajectory_focus}_litc_{decompose_method}.pkl")
  859. with open(input_trajectories_path_litc, 'rb') as _file:
  860. lit = pickle.load(_file)
  861. break
  862. df_edges = []
  863. for ind, row in result_df.iterrows():
  864. trj = lit.get_trajectory(row["booted_trajectory"], False)
  865. for from_cell, to_cell in trj.edges():
  866. row_copy = row.copy()
  867. row_copy["from_cell"] = from_cell
  868. row_copy["to_cell"] = to_cell
  869. df_edges.append(row_copy)
  870. df_edges = pd.DataFrame(df_edges)
  871. df_edges.reset_index(drop=True, inplace=True)
  872. df_edges_pivot = df_edges.pivot(index=['path', 'booted_trajectory', 'trajectory', 'metric', 'from_cell', 'to_cell'], columns='representation', values='score')
  873. df_edges_pivot_rank = rank_metrics(df_edges_pivot, metric_direction_dict, metric_index_pos=3, method='min')
  874. df_edges_pivot_rank_mean = logistic_invert_rowwise(df_edges_pivot_rank).groupby(level=[4, 5]).mean()
  875. df_edges_pivot_rank_p = logistic_invert_rowwise(df_edges_pivot_rank).groupby(['from_cell', 'to_cell']).apply(compare_scores_and_get_means).reset_index()
  876. alpha = 0.05 / len(df_edges_pivot_rank_p)
  877. df_edges_pivot_rank_p['mean_scvi_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scvi'], np.nan)
  878. df_edges_pivot_rank_p['mean_tardis_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_tardis'], np.nan)
  879. trj = InputTrajectories(ground_truth_trajectories).get_trajectory(trajectory_focus, False)
  880. representations = ["scvi", "tardis"]
  881. for ind, row in df_edges_pivot_rank_p.iterrows():
  882. for repr in representations:
  883. trj[row["from_cell"]][row["to_cell"]][f"weight_{repr}"] = row[f"mean_{repr}_final"]
  884. # %%
  885. for i in ['weight_tardis', 'weight_scvi']:
  886. plot_trajectory(
  887. trj,
  888. figsize=(4, 3.25),
  889. edge_weight_attribute=i,
  890. title="",
  891. edge_vmin=0.4,
  892. edge_vmax=0.7,
  893. color_edges_by_weight=True,
  894. node_size=150,
  895. node_color="lightgray",
  896. edge_color='black',
  897. edge_width=5,
  898. edge_labels=False,
  899. edge_label_font_size=5,
  900. edge_cmap='viridis',
  901. edge_colorbar=False,
  902. file_name=f"{trajectory_focus}_{i}"
  903. )
  904. # %%
  905. from matplotlib import ticker
  906. def create_vertical_colorbar(
  907. label,
  908. filename,
  909. vmin=2.5,
  910. vmax=3.5,
  911. cmap='plasma',
  912. fontsize=10,
  913. n_ticks=5,
  914. fig_height=4,
  915. fig_width=1,
  916. scale=1.0
  917. ):
  918. """
  919. Creates and displays a standalone vertical colorbar.
  920. Parameters:
  921. - label: str, label for the colorbar
  922. - vmin, vmax: float, data range
  923. - cmap: str, colormap name
  924. - fontsize: int, for label & tick labels
  925. - n_ticks: int, approx number of ticks to show
  926. - fig_height: float, base height in inches
  927. - fig_width: float, base width in inches
  928. - scale: float, scale factor for overall size
  929. """
  930. # 1) Make the ScalarMappable
  931. norm = mpl.colors.Normalize(vmin=vmin, vmax=vmax)
  932. sm = mpl.cm.ScalarMappable(norm=norm, cmap=cmap)
  933. sm.set_array([])
  934. # 2) New figure, scaled
  935. fig = plt.figure(figsize=(fig_width * scale, fig_height * scale))
  936. # you can tweak these numbers if you want more/less padding
  937. cax = fig.add_axes([0.2, 0.05, 0.1, 0.9])
  938. # 3) Draw the colorbar
  939. cbar = fig.colorbar(
  940. sm,
  941. cax=cax,
  942. orientation='vertical'
  943. )
  944. # 4) Automatic “nice” locators + re‐draw
  945. cbar.locator = ticker.MaxNLocator(nbins=n_ticks)
  946. cbar.update_ticks()
  947. # 5) Force one‐decimal formatting
  948. cbar.ax.yaxis.set_major_formatter(ticker.FormatStrFormatter('%.1f'))
  949. # 6) Label & style
  950. cbar.set_label(label, fontsize=fontsize)
  951. cbar.ax.tick_params(labelsize=fontsize)
  952. plt.subplots_adjust(left=0.03, right=0.97, top=0.97, bottom=0.15)
  953. plt.savefig(f"/Users/kemalinecik/Downloads/fig2fig/{filename}.pdf")
  954. create_vertical_colorbar("Performance", filename = "colorbar_performance_rank",
  955. cmap="viridis", n_ticks=3, fig_height=2, fig_width=2.4, scale=0.8, fontsize=10, vmin=0.4, vmax=0.7)
  956. # %% [markdown]
  957. # # Figure 3
  958. # %%
  959. result_df_boot = pd.read_pickle(os.path.join(temp_move, f"df_boots_fig2_calculations.pkl"))
  960. result_df_boot = result_df_boot[result_df_boot["representation"]=="scvi"]
  961. result_df_boot.reset_index(inplace=True, drop=True)
  962. result_df_main = pd.read_pickle(os.path.join(temp_move, f"df_boots_fig2_calculations_for_individuals.pkl"))
  963. # %%
  964. focus_trajectory = "germline"
  965. df_compo = result_df_main[result_df_main["trajectory"] == focus_trajectory]
  966. df_atlas = result_df_boot[result_df_boot["trajectory"] == focus_trajectory]
  967. df_atlas.reset_index(inplace=True, drop=True)
  968. df_compo.reset_index(inplace=True, drop=True)
  969. decompose_method = "2"
  970. input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{focus_trajectory}_litc_{decompose_method}.pkl")
  971. with open(input_trajectories_path_litc, 'rb') as _file:
  972. lit = pickle.load(_file)
  973. trj_size_dict = dict()
  974. for trj_name in lit.graph['trajectories']:
  975. trj_size_dict[trj_name] = len(lit.get_trajectory(trj_name, False).nodes()), len(lit.get_trajectory(trj_name, False).edges())
  976. assert df_compo.shape == df_atlas.shape
  977. for i in range(len(df_compo)):
  978. i1 = list(df_atlas.iloc[i][["booted_trajectory", "metric"]])
  979. i2 = list(df_compo.iloc[i][["booted_trajectory", "metric"]])
  980. assert i1==i2
  981. # df_both
  982. df_both = df_compo.copy()
  983. del df_both["score"]
  984. df_both["score_Component"] = list(df_compo["score"])
  985. df_both["score_Atlas"] = list(df_atlas["score"])
  986. df_both.reset_index(inplace=True, drop=True)
  987. del df_both["anndata_subset"]
  988. def get_winners(row):
  989. metric = row['metric']
  990. direction = metric_direction_dict[metric]
  991. c = row['score_Component']
  992. a = row['score_Atlas']
  993. if direction == 'decreasing':
  994. if c > a:
  995. return pd.Series([0, 1])
  996. elif a > c:
  997. return pd.Series([1, 0])
  998. else:
  999. return pd.Series([0.5, 0.5])
  1000. elif direction == 'increasing':
  1001. if c < a:
  1002. return pd.Series([0, 1])
  1003. elif a < c:
  1004. return pd.Series([1, 0])
  1005. else:
  1006. return pd.Series([0.5, 0.5])
  1007. else:
  1008. return pd.Series([None, None]) # In case metric not found
  1009. df_both[['winner_Component', 'winner_Atlas']] = df_both.apply(get_winners, axis=1)
  1010. def get_nodes_edges(row):
  1011. trj_name = row['booted_trajectory']
  1012. len_nodes, len_edges = trj_size_dict[trj_name]
  1013. return pd.Series([len_nodes, len_edges])
  1014. df_both[['len_nodes', 'len_edges']] = df_both.apply(get_nodes_edges, axis=1)
  1015. label_mapping_o = {
  1016. "adjacency": "Adjacency",
  1017. "embedding": "Embedding",
  1018. "pseudotime": "Pseudotime",
  1019. }
  1020. df_both["path"] = [label_mapping_o[i] for i in df_both["path"]]
  1021. # long_both
  1022. id_cols = [
  1023. "decompose_method", "trajectory_class", "trajectory",
  1024. "booted_trajectory", 'len_nodes', 'len_edges', "path", "metric",
  1025. "trajectory_annotation_column"
  1026. ]
  1027. _df = df_both.reset_index()
  1028. id_cols = ["index"] + id_cols
  1029. long_both = (
  1030. pd.wide_to_long(
  1031. _df, # the original wide dataframe
  1032. stubnames=["score", "winner"],
  1033. i=id_cols, # columns that make a unique row
  1034. j="focus", # new column: 'component' or 'atlas'
  1035. sep="_", # underscore separates stub & suffix
  1036. suffix="(Component|Atlas)" # the two suffix values to capture
  1037. )
  1038. .reset_index()
  1039. .sort_values(id_cols + ["focus"]) # optional tidy ordering
  1040. )
  1041. del long_both["index"]
  1042. replacements = {
  1043. 'Pseudotime': 'Cell Order\ncontinuity',
  1044. 'Embedding': 'Trajectory\ndominance',
  1045. 'Adjacency': 'Topology\nconservation'
  1046. }
  1047. long_both_p = long_both.pivot(index=["booted_trajectory", "path", "metric"], columns='focus', values='winner')
  1048. long_both_p = long_both_p.groupby(level=[0, 1]).mean()
  1049. long_both_p.rename(index=replacements, level="path", inplace=True)
  1050. long_both_back = (
  1051. long_both_p
  1052. .stack(dropna=False) # <- dropna=True by default
  1053. .rename('winner')
  1054. .reset_index() # gives: booted_trajectory, path, metric, focus, winner
  1055. )
  1056. long_both_back.head()
  1057. # %%
  1058. label_mapping = {
  1059. "Atlas": "IA",
  1060. "Component": "Garcia",
  1061. }
  1062. # match boxplot’s height & per-column width
  1063. fig_height = 3
  1064. n_cols = long_both_back["focus"].nunique()
  1065. col_width = 2.4
  1066. fig_width = n_cols * col_width
  1067. fig, ax = plt.subplots(figsize=(fig_width, fig_height))
  1068. # draw bars
  1069. bp = sns.barplot(
  1070. data=long_both_back,
  1071. hue="focus",
  1072. y="winner",
  1073. x="path",
  1074. palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
  1075. width=0.8,
  1076. saturation=1,
  1077. errorbar="sd",
  1078. capsize=0.05,
  1079. err_kws={'color': 'black', 'linewidth': 0.81},
  1080. edgecolor="black",
  1081. linewidth=0.81,
  1082. zorder=1,
  1083. ax=ax
  1084. )
  1085. # enforce edge styling
  1086. for patch in bp.patches:
  1087. patch.set_edgecolor("black")
  1088. patch.set_linewidth(0.81)
  1089. # axes
  1090. ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
  1091. ax.tick_params(axis="y", labelsize=10)
  1092. ax.set_xlabel('Germline', fontsize=10, labelpad=10)
  1093. plt.ylabel("Performance", fontsize=10, labelpad=10)
  1094. ax.set_ylim(0, None)
  1095. ax.yaxis.set_major_formatter(FormatStrFormatter('%.1f'))
  1096. # grid + despine
  1097. ax.grid(axis='y', linestyle='--', alpha=0.4)
  1098. sns.despine(left=True, trim=False)
  1099. # remove the old legend
  1100. ax.legend_.remove()
  1101. # build new, professional legend
  1102. handles, labels = ax.get_legend_handles_labels()
  1103. new_labels = [label_mapping.get(l, l) for l in labels]
  1104. slim_handles = [
  1105. mpatches.Patch(
  1106. facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
  1107. edgecolor="black",
  1108. linewidth=0.5 # thinner border
  1109. )
  1110. for h in handles
  1111. ]
  1112. leg = ax.legend(
  1113. slim_handles,
  1114. new_labels,
  1115. title="Representation",
  1116. loc="center left",
  1117. bbox_to_anchor=(1.02, 0.5),
  1118. alignment="left",
  1119. frameon=False,
  1120. title_fontsize=8,
  1121. fontsize=8,
  1122. borderaxespad=0.5,
  1123. handletextpad=0.4,
  1124. labelspacing=0.5,
  1125. )
  1126. # tighten layout (accounts for legend)
  1127. plt.tight_layout(rect=[0, 0, 0.85, 1])
  1128. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/comparison_germline_component.pdf")
  1129. # %%
  1130. long_both_germ = long_both_back.copy()
  1131. # %%
  1132. focus_trajectory = "Haem_ref"
  1133. df_compo = result_df_main[result_df_main["trajectory"] == focus_trajectory]
  1134. df_atlas = result_df_boot[result_df_boot["trajectory"] == focus_trajectory]
  1135. df_atlas.reset_index(inplace=True, drop=True)
  1136. df_compo.reset_index(inplace=True, drop=True)
  1137. decompose_method = "2"
  1138. input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{focus_trajectory}_litc_{decompose_method}.pkl")
  1139. with open(input_trajectories_path_litc, 'rb') as _file:
  1140. lit = pickle.load(_file)
  1141. trj_size_dict = dict()
  1142. for trj_name in lit.graph['trajectories']:
  1143. trj_size_dict[trj_name] = len(lit.get_trajectory(trj_name, False).nodes()), len(lit.get_trajectory(trj_name, False).edges())
  1144. assert df_compo.shape == df_atlas.shape
  1145. for i in range(len(df_compo)):
  1146. i1 = list(df_atlas.iloc[i][["booted_trajectory", "metric"]])
  1147. i2 = list(df_compo.iloc[i][["booted_trajectory", "metric"]])
  1148. assert i1==i2
  1149. # df_both
  1150. df_both = df_compo.copy()
  1151. del df_both["score"]
  1152. df_both["score_Component"] = list(df_compo["score"])
  1153. df_both["score_Atlas"] = list(df_atlas["score"])
  1154. df_both.reset_index(inplace=True, drop=True)
  1155. del df_both["anndata_subset"]
  1156. def get_winners(row):
  1157. metric = row['metric']
  1158. direction = metric_direction_dict[metric]
  1159. c = row['score_Component']
  1160. a = row['score_Atlas']
  1161. if direction == 'decreasing':
  1162. if c > a:
  1163. return pd.Series([0, 1])
  1164. elif a > c:
  1165. return pd.Series([1, 0])
  1166. else:
  1167. return pd.Series([0.5, 0.5])
  1168. elif direction == 'increasing':
  1169. if c < a:
  1170. return pd.Series([0, 1])
  1171. elif a < c:
  1172. return pd.Series([1, 0])
  1173. else:
  1174. return pd.Series([0.5, 0.5])
  1175. else:
  1176. return pd.Series([None, None]) # In case metric not found
  1177. df_both[['winner_Component', 'winner_Atlas']] = df_both.apply(get_winners, axis=1)
  1178. def get_nodes_edges(row):
  1179. trj_name = row['booted_trajectory']
  1180. len_nodes, len_edges = trj_size_dict[trj_name]
  1181. return pd.Series([len_nodes, len_edges])
  1182. df_both[['len_nodes', 'len_edges']] = df_both.apply(get_nodes_edges, axis=1)
  1183. label_mapping_o = {
  1184. "adjacency": "Adjacency",
  1185. "embedding": "Embedding",
  1186. "pseudotime": "Pseudotime",
  1187. }
  1188. df_both["path"] = [label_mapping_o[i] for i in df_both["path"]]
  1189. # long_both
  1190. id_cols = [
  1191. "decompose_method", "trajectory_class", "trajectory",
  1192. "booted_trajectory", 'len_nodes', 'len_edges', "path", "metric",
  1193. "trajectory_annotation_column"
  1194. ]
  1195. _df = df_both.reset_index()
  1196. id_cols = ["index"] + id_cols
  1197. long_both = (
  1198. pd.wide_to_long(
  1199. _df, # the original wide dataframe
  1200. stubnames=["score", "winner"],
  1201. i=id_cols, # columns that make a unique row
  1202. j="focus", # new column: 'component' or 'atlas'
  1203. sep="_", # underscore separates stub & suffix
  1204. suffix="(Component|Atlas)" # the two suffix values to capture
  1205. )
  1206. .reset_index()
  1207. .sort_values(id_cols + ["focus"]) # optional tidy ordering
  1208. )
  1209. del long_both["index"]
  1210. replacements = {
  1211. 'Pseudotime': 'Cell Order\ncontinuity',
  1212. 'Embedding': 'Trajectory\ndominance',
  1213. 'Adjacency': 'Topology\nconservation'
  1214. }
  1215. long_both_p = long_both.pivot(index=["booted_trajectory", "path", "metric"], columns='focus', values='winner')
  1216. long_both_p = long_both_p.groupby(level=[0, 1]).mean()
  1217. long_both_p.rename(index=replacements, level="path", inplace=True)
  1218. long_both_back = (
  1219. long_both_p
  1220. .stack(dropna=False) # <- dropna=True by default
  1221. .rename('winner')
  1222. .reset_index() # gives: booted_trajectory, path, metric, focus, winner
  1223. )
  1224. long_both_back.head()
  1225. # %%
  1226. label_mapping = {
  1227. "Atlas": "IA",
  1228. "Component": "Suo",
  1229. }
  1230. # match boxplot’s height & per-column width
  1231. fig_height = 3
  1232. n_cols = long_both_back["focus"].nunique()
  1233. col_width = 2.4
  1234. fig_width = n_cols * col_width
  1235. fig, ax = plt.subplots(figsize=(fig_width, fig_height))
  1236. # draw bars
  1237. bp = sns.barplot(
  1238. data=long_both_back,
  1239. hue="focus",
  1240. y="winner",
  1241. x="path",
  1242. palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
  1243. width=0.8,
  1244. saturation=1,
  1245. errorbar="sd",
  1246. capsize=0.05,
  1247. err_kws={'color': 'black', 'linewidth': 0.81},
  1248. edgecolor="black",
  1249. linewidth=0.81,
  1250. zorder=1,
  1251. ax=ax
  1252. )
  1253. # enforce edge styling
  1254. for patch in bp.patches:
  1255. patch.set_edgecolor("black")
  1256. patch.set_linewidth(0.81)
  1257. # axes
  1258. ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
  1259. ax.tick_params(axis="y", labelsize=10)
  1260. ax.set_xlabel('Hematopoietic Lineage', fontsize=10, labelpad=10)
  1261. plt.ylabel("Performance", fontsize=10, labelpad=10)
  1262. ax.set_ylim(0, None)
  1263. ax.yaxis.set_major_formatter(FormatStrFormatter('%.1f'))
  1264. # grid + despine
  1265. ax.grid(axis='y', linestyle='--', alpha=0.4)
  1266. sns.despine(left=True, trim=False)
  1267. # remove the old legend
  1268. ax.legend_.remove()
  1269. # build new, professional legend
  1270. handles, labels = ax.get_legend_handles_labels()
  1271. new_labels = [label_mapping.get(l, l) for l in labels]
  1272. slim_handles = [
  1273. mpatches.Patch(
  1274. facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
  1275. edgecolor="black",
  1276. linewidth=0.5 # thinner border
  1277. )
  1278. for h in handles
  1279. ]
  1280. leg = ax.legend(
  1281. slim_handles,
  1282. new_labels,
  1283. title="Representation",
  1284. loc="center left",
  1285. bbox_to_anchor=(1.02, 0.5),
  1286. alignment="left",
  1287. frameon=False,
  1288. title_fontsize=8,
  1289. fontsize=8,
  1290. borderaxespad=0.5,
  1291. handletextpad=0.4,
  1292. labelspacing=0.5,
  1293. )
  1294. # tighten layout (accounts for legend)
  1295. plt.tight_layout(rect=[0, 0, 0.85, 1])
  1296. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/comparison_hemato_component.pdf")
  1297. # %%
  1298. long_both_haem = long_both_back.copy()
  1299. # %%
  1300. long_both_haem["trajectory_class"] = "Hematopoietic\nLineage"
  1301. long_both_germ["trajectory_class"] = "Germline"
  1302. long_both2 = pd.concat([long_both_haem, long_both_germ])
  1303. long_both2.reset_index(drop=True, inplace=True)
  1304. # %%
  1305. label_mapping = {
  1306. "Atlas": "Atlas",
  1307. "Component": "Component",
  1308. }
  1309. # match boxplot’s height & per-column width
  1310. fig_height = 3
  1311. n_cols = long_both2["focus"].nunique()
  1312. col_width = 2
  1313. fig_width = n_cols * col_width
  1314. fig, ax = plt.subplots(figsize=(fig_width, fig_height))
  1315. # draw bars
  1316. bp = sns.barplot(
  1317. data=long_both2,
  1318. hue="focus",
  1319. y="winner",
  1320. x="trajectory_class",
  1321. palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
  1322. width=0.8,
  1323. saturation=1,
  1324. errorbar="sd",
  1325. capsize=0.05,
  1326. err_kws={'color': 'black', 'linewidth': 0.81},
  1327. edgecolor="black",
  1328. linewidth=0.81,
  1329. zorder=1,
  1330. ax=ax
  1331. )
  1332. # enforce edge styling
  1333. for patch in bp.patches:
  1334. patch.set_edgecolor("black")
  1335. patch.set_linewidth(0.81)
  1336. # axes
  1337. ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
  1338. ax.tick_params(axis="y", labelsize=10)
  1339. ax.set_xlabel('Lineage', fontsize=10, labelpad=10)
  1340. plt.ylabel("Performance", fontsize=10, labelpad=10)
  1341. ax.set_ylim(0, None)
  1342. ax.yaxis.set_major_formatter(FormatStrFormatter('%.1f'))
  1343. # grid + despine
  1344. ax.grid(axis='y', linestyle='--', alpha=0.4)
  1345. sns.despine(left=True, trim=False)
  1346. # remove the old legend
  1347. ax.legend_.remove()
  1348. # build new, professional legend
  1349. handles, labels = ax.get_legend_handles_labels()
  1350. new_labels = [label_mapping.get(l, l) for l in labels]
  1351. slim_handles = [
  1352. mpatches.Patch(
  1353. facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
  1354. edgecolor="black",
  1355. linewidth=0.5 # thinner border
  1356. )
  1357. for h in handles
  1358. ]
  1359. leg = ax.legend(
  1360. slim_handles,
  1361. new_labels,
  1362. title="Representation",
  1363. loc="center left",
  1364. bbox_to_anchor=(1.02, 0.5),
  1365. alignment="left",
  1366. frameon=False,
  1367. title_fontsize=8,
  1368. fontsize=8,
  1369. borderaxespad=0.5,
  1370. handletextpad=0.4,
  1371. labelspacing=0.5,
  1372. )
  1373. # tighten layout (accounts for legend)
  1374. plt.tight_layout(rect=[0, 0, 0.85, 1])
  1375. plt.savefig("/Users/kemalinecik/Downloads/fig2fig/comparison_germline_hemato_both.pdf")
  1376. # %%
  1377. 1
  1378. # %% [markdown]
  1379. # # Figure UMAPs
  1380. # %%
  1381. import scvelo as scv
  1382. import cellrank as cr
  1383. import scanpy.external as sce
  1384. scv.settings.verbosity = 3
  1385. cr.settings.verbosity = 2
  1386. sc.settings.verbose = 3
  1387. # %%
  1388. adata_comp = ad.read_h5ad(os.path.join(temp_move, f"anndata_hdca_input_germline_litc_comparison_comp.h5ad"))
  1389. adata_atlas = ad.read_h5ad(os.path.join(temp_move, f"anndata_hdca_input_germline_litc_comparison_atlas.h5ad"))
  1390. # %%
  1391. display((adata_atlas.shape, adata_comp.shape))
  1392. sc.pp.neighbors(adata_atlas, n_neighbors=15)
  1393. sc.tl.umap(adata_atlas)
  1394. sc.pp.neighbors(adata_comp, n_neighbors=15)
  1395. sc.tl.umap(adata_comp)
  1396. # %%
  1397. (
  1398. df_hdca_component_adata_paths,
  1399. df_hdca_lvl3_adata_paths,
  1400. out_df,
  1401. ground_truth_trajectories,
  1402. ground_truth_trajectories_annotation_level,
  1403. trajectory_subset_dict
  1404. ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
  1405. # %%
  1406. # Mapping from original to formal names
  1407. celltype_mapping = {
  1408. "GC": "Granulosa Cell",
  1409. "PGC": "Primordial Germ Cell",
  1410. "oocyte": "Oocyte",
  1411. "oogonia": "Oogonia",
  1412. "oogonia_meiotic_1": "Oogonia\n(Meiotic I)",
  1413. "oogonia_meiotic_2": "Oogonia\n(Meiotic II)",
  1414. "pre_oocyte": "Pre-Oocyte",
  1415. "pre_spermatogonia": "Pre-Spermatogonia"
  1416. }
  1417. # Apply mapping to the annotations
  1418. adata_atlas.obs[ground_truth_trajectories_annotation_level[focus_trajectory]] = \
  1419. adata_atlas.obs[ground_truth_trajectories_annotation_level[focus_trajectory]].map(celltype_mapping)
  1420. # Apply mapping to the annotations
  1421. adata_comp.obs[ground_truth_trajectories_annotation_level[focus_trajectory]] = \
  1422. adata_comp.obs[ground_truth_trajectories_annotation_level[focus_trajectory]].map(celltype_mapping)
  1423. # %%
  1424. def cr_calc(adata_cr, root_cell_type, i, ann_levl, knn=15):
  1425. sc.tl.pca(adata_cr, n_comps=11, svd_solver='arpack')
  1426. sce.tl.palantir(adata_cr, n_components=10, knn=knn)
  1427. root_cells = adata_cr.obs_names[adata_cr.obs[ann_levl] == root_cell_type]
  1428. if len(root_cells) == 0:
  1429. raise ValueError(f"No cells found with {ann_levl!r} == {root_cell_type!r}")
  1430. start_cell = root_cells[i]
  1431. pr_res = sce.tl.palantir_results(
  1432. adata_cr,
  1433. early_cell=start_cell,
  1434. ms_data="X_palantir_multiscale",
  1435. knn=knn,
  1436. num_waypoints=1200,
  1437. n_jobs=-1
  1438. )
  1439. adata_cr.obs["palantir_pseudotime"] = pr_res.pseudotime
  1440. pk = cr.kernels.PseudotimeKernel(adata_cr, time_key='palantir_pseudotime')
  1441. pk.compute_transition_matrix()
  1442. return pk
  1443. # %%
  1444. focus_trajectory = "germline"
  1445. ann_lvl = ground_truth_trajectories_annotation_level[focus_trajectory]
  1446. pk_adata_atlas = cr_calc(adata_atlas.copy(), root_cell_type = "Primordial Germ Cell", i=10, ann_levl=ann_lvl)
  1447. # %%
  1448. plt.figure(figsize=(4,3))
  1449. pk_adata_atlas.plot_projection(basis='umap', recompute=False, color=[ann_lvl], title="", ax=plt.gca(), legend_fontsize="medium", legend_fontweight="bold", legend_fontoutline=2,
  1450. save="/Users/kemalinecik/Downloads/fig2fig/comparison_germline_atlas_umap.pdf")
  1451. # %%
  1452. pk_adata_comp = cr_calc(adata_comp.copy(), root_cell_type = "Primordial Germ Cell", i=10, ann_levl=ann_lvl)
  1453. # %%
  1454. plt.figure(figsize=(4,3))
  1455. pk_adata_comp.plot_projection(basis='umap', recompute=False, color=[ann_lvl], title="", ax=plt.gca(), legend_fontsize="medium", legend_fontweight="bold", legend_fontoutline=2,
  1456. save="/Users/kemalinecik/Downloads/fig2fig/comparison_germline_component_umap.pdf")
  1457. # %% [markdown]
  1458. # # Figure UMAPs
  1459. # %%
  1460. import scvelo as scv
  1461. import cellrank as cr
  1462. import scanpy.external as sce
  1463. scv.settings.verbosity = 3
  1464. cr.settings.verbosity = 2
  1465. sc.settings.verbose = 3
  1466. (
  1467. df_hdca_component_adata_paths,
  1468. df_hdca_lvl3_adata_paths,
  1469. out_df,
  1470. ground_truth_trajectories,
  1471. ground_truth_trajectories_annotation_level,
  1472. trajectory_subset_dict
  1473. ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
  1474. # %%
  1475. def cr_calc(adata_cr, root_cell_type, i, ann_levl, knn=15):
  1476. sc.tl.pca(adata_cr, n_comps=11, svd_solver='arpack')
  1477. sce.tl.palantir(adata_cr, n_components=10, knn=knn)
  1478. root_cells = adata_cr.obs_names[adata_cr.obs[ann_levl] == root_cell_type]
  1479. if len(root_cells) == 0:
  1480. raise ValueError(f"No cells found with {ann_levl!r} == {root_cell_type!r}")
  1481. start_cell = root_cells[i]
  1482. pr_res = sce.tl.palantir_results(
  1483. adata_cr,
  1484. early_cell=start_cell,
  1485. ms_data="X_palantir_multiscale",
  1486. knn=knn,
  1487. num_waypoints=1200,
  1488. n_jobs=-1
  1489. )
  1490. adata_cr.obs["palantir_pseudotime"] = pr_res.pseudotime
  1491. pk = cr.kernels.PseudotimeKernel(adata_cr, time_key='palantir_pseudotime')
  1492. pk.compute_transition_matrix()
  1493. return pk
  1494. # %%
  1495. # Mapping for B-cell and progenitor annotations
  1496. bcell_mapping = {
  1497. "Mature_B": "Mature B",
  1498. "Large_pre_B": "Large Pre B",
  1499. "Small_pre_B": "Small Pre B",
  1500. "B1": "B1",
  1501. "Pro_B": "Pro B",
  1502. "Late_Pro_B": "Late Pro B",
  1503. "HSC_MPP": "HSC MPP",
  1504. "Immature_B": "Immature B",
  1505. "Pre_pro_B": "Pre-pro B",
  1506. "LMPP_MLP": "LMPP MLP",
  1507. "CYCLING_B": "Cycling B",
  1508. "Plasma_B": "Plasma B"
  1509. }
  1510. # %%
  1511. for use_rep in ['scpoli', 'scvi']:
  1512. print("\n\n\n\%%%%%%%%%%%%%%\n\n\n", use_rep)
  1513. adata_atlas = ad.read_h5ad(os.path.join(temp_move, f"anndata_hdca_input_Haem_ref_litc_2_{use_rep}_nodes_b.h5ad"))
  1514. print(adata_atlas.shape)
  1515. sc.pp.subsample(adata_atlas, fraction=0.2)
  1516. print(adata_atlas.shape)
  1517. sc.pp.neighbors(adata_atlas, n_neighbors=15)
  1518. sc.tl.umap(adata_atlas)
  1519. ann_lvl = ground_truth_trajectories_annotation_level["Haem_ref"]
  1520. adata_atlas.obs[ann_lvl] = adata_atlas.obs[ann_lvl].map(bcell_mapping)
  1521. pk_adata_atlas = cr_calc(adata_atlas.copy(), root_cell_type = "HSC MPP", i=10, ann_levl=ann_lvl, knn=30)
  1522. plt.figure(figsize=(5,4))
  1523. pk_adata_atlas.plot_projection(basis='umap', recompute=False, color=[ann_lvl], title="", ax=plt.gca(), legend_fontsize="medium", legend_fontweight="bold", legend_fontoutline=2,
  1524. save=f"/Users/kemalinecik/Downloads/fig2fig/comparison_models_b_umap_{use_rep}.svg")
  1525. # break
  1526. del adata_atlas, pk_adata_atlas
  1527. gc.collect()
  1528. # %%

sctram_figures_only.ipynb at commit e570de5, under BSD-3-Clause · at the source

Overview

Authors: Simone Webb1,2, Antony Rose1,2, Chuan Xu3, Lloyd Steele2, Masami A Kuri2, Emily Stephenson1,2, Kemal Inecik4, Daniyal Jafree2,5, April R Foster2, Daniela Basurto-Lozada1,2, Nana-Jane Chipampe2, Anna V Pournara2, Marc-Antoine Jacques6,7, Ken To2,3,8, Chloe Admane1,2, Efpraxia Kritikaki1,2, Marta A Chroscik1,2, David Horsfall1,2, Julia Foreman6, Koen Rademaker2
and 74 other authorsJulia Karjalainen9, Anna Laddach10, Shaista Madad2, John EG Lawrence2,8, Vitalii Kleshchevnikov2, Steven Lisgo1, Jimmy Tsz Hang Lee2, Joséphine Blevinal11, Ahlam Alqahtani1, Stanislaw Makarchuk2, Jake Jackson1, Ekin Ucuncu5, Teresa P Silva5, Valentina Lorenzi2,6, Fereshteh Torabi2, Rachel A Botting1, Kenny Roberts2, Bayanne Olabi1, Keerthi P Chakala2, Leander Dony4,12, Gabriele Dall’Aglio4, Ana-Maria Cujba3, Holly J Whitfield2, Christine Hale2, Justin Engelbert1, Tong Li2, Elena Prigmore2, Nusayhah Hudaa Gopee1,2, Michael W Mather1, Ben Talks1, Lijiang Fei3, Diana Pereira13, Corina Moldovan14, Rasa Elmentaite3, Krzysztof Polanski3, Alexander V Predeus2, Pasha Mazin2, Vicky Rowe2, Semih Bayraktar3, Vincent R Knight-Schrijver3, Menna Clatworthy15, Kazumasa Kanemaru3, James Cranley3,16, Andrew Guo15, Sanjay Sinha3, David McDonald1, Andrew Filby1, Janet Kerwin1, Peng He17,18, Inês Sequeira13, Roser Vento-Tormo2, Malte D. Luecken4,19, Alain Chédotal11,20,21, Deanne Taylor22,23, Silvia Martin-Almedina24, Pia Ostergaard24, Sam Behjati2,25, Elizabeth Robertson26, Nicola K Wilson3,7, John Marioni5,27, Helen V Firth2,28, Caroline F Wright29, Olivier Pourquie30, Laure Gambardella2, Paula Alexandre5, Berthold Gottgens3,7, Kerstin B Meyer2, MS Vijayabaskar2, Mo Lotfollahi2, Sarah A Teichmann3,16,31, Fabian J Theis4, Fränze Progatzky9, Omer Ali Bayraktar2, Muzlifah Haniffa2,28,32
32 affiliations
  1. Biosciences Institute, Newcastle University, Newcastle upon Tyne, UK
  2. Wellcome Sanger Institute, Wellcome Genome Campus, Hinxton, UK
  3. Cambridge Stem Cell Institute, Jeffrey Cheah Biomedical Centre, Cambridge Biomedical Campus, University of Cambridge, Cambridge, UK
  4. Institute of Computational Biology, Computational Health Center, Helmholtz Munich, Neuherberg, Germany
  5. Developmental Biology & Cancer Department, University College London (UCL) Great Ormond Street Institute of Child Health, UCL, London, UK
  6. European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Cambridge, UK
  7. Department of Haematology, University of Cambridge, Cambridge, UK
  8. Department of Surgery, University of Cambridge, Cambridge, UK
  9. Kennedy Institute of Rheumatology, University of Oxford, Oxford, UK
  10. Nervous System Development and Homeostasis Laboratory, the Francis Crick Institute, London, UK
  11. Sorbonne Université, INSERM, CNRS, Institut de la Vision, Paris, France
  12. Department of Biosystems Science and Engineering, ETH Zürich, Basel, Switzerland
  13. Center for Oral Immunobiology and Regenerative Medicine, Barts Centre for Squamous Cancer, Institute of Dentistry, Queen Mary University of London, London, UK
  14. Department of Pathology, Newcastle Hospitals NHS Foundation Trust, Newcastle upon Tyne, UK
  15. Molecular Immunity Unit, Department of Medicine, University of Cambridge, Cambridge, UK
  16. Department of Medicine, University of Cambridge, Cambridge, UK
  17. Department of Pathology, University of California, San Francisco, CA, USA
  18. Eli and Edythe Broad Center of Regeneration Medicine and Stem Cell Research, University of California, San Francisco, San Francisco, CA, USA
  19. Institute of Lung Health and Immunity (LHI), Helmholtz Munich, Comprehensive Pneumology Center (CPC-M), Germany; Member of the German Center for Lung Research (DZL)
  20. Institut de Pathologie, Groupe Hospitalier Est, Hospices Civils de Lyon, Lyon, France
  21. University Claude Bernard Lyon 1, MeLiS, CNRS UMR5284, INSERM U1314, 69008, Lyon, France
  22. Department of Biomedical and Health Informatics (DBHi), The Children’s Hospital of Philadelphia, Philadelphia, PA, USA
  23. Department of Pediatrics, University of Pennsylvania Perelman School of Medicine, Philadelphia, PA, USA
  24. School of Health and Medical Sciences, City St George’s University of London, London, UK
  25. Department of Paediatrics, University of Cambridge, Cambridge, UK
  26. Sir William Dunn School of Pathology, University of Oxford, Oxford, UK
  27. Cancer Research UK, Cambridge Institute, University of Cambridge, Cambridge, UK
  28. Cambridge University Hospitals Foundation Trust, Addenbrooke’s Hospital, Cambridge, UK
  29. Department of Clinical and Biomedical Sciences, St Luke’s Campus, University of Exeter, Exeter, UK
  30. Department of Genetics, Harvard Medical School and Department of Pathology, Brigham and Women’s Hospital, Boston, MA, USA
  31. CIFAR Macmillan Multi-scale Human Programme, CIFAR, Toronto, Canada
  32. Cambridge Institute for Therapeutic Immunology and Infectious Diseases (CIITID), University of Cambridge, UK
Dates: published online 1 April 2026
Type: Preprint
License: CC BY
Identifiers: DOI 10.64898/2026.03.30.714220 · OpenAlex W7147094455
Open access: green, a free copy (OpenAlex)
Status: code verified
Categories: human (organism)
Methods: Preprocessing, Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, fMRI & imaging
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Institute for Health Research (NIHR) (NIHR203312); Wellcome Trust (WT223718/Z/21/Z, 203151/Z/16/Z, 226083/Z/22/Z, 203151/A/16/Z, 314710/Z/24/Z); British Heart Foundation (RG/17/7/33217, SP/13/5/30288); Biotechnology and Biological Sciences Research Council (BB/Y003179/1); Versus Arthritis (22981)
Citations: cited by 2 papers (Europe PMC); 193 references in the paper

Abstract

Single cell genomics has enabled analysis of human prenatal development at unprecedented resolution. However, most studies have relied on dissociated tissues during restricted windows of development, limiting insights into how spatially distributed networks of cells, and multicellular niches emerge and adapt to distinct organ microenvironments in situ. Moreover, existing human developmental atlases have not yet been harmonised, and we thus lack a comprehensive catalogue of known cell types in the developing human body.

Here, we introduce the Human Developmental Cell Atlas (HDCA), a unified structural, cellular and molecular resource for prenatal human development. The HDCA integrates published and unpublished single cell/nucleus RNAseq atlases across prenatal organs, and includes a newly generated, spatially resolved, multimodal cell atlas of intact human embryos. Spanning 4-22 post conceptional weeks, capturing embryonic and early to mid fetal stages, the HDCA contains ~4.6 million cells/nuclei which resolve into ~450 cell types, explorable with a bespoke web portal.

For a global overview of the human embryo’s multicellular communities, we applied unsupervised deep learning to our intact human embryo spatial data, charting 114 tissue niches that are structural and signalling hubs for the cellular interactions of the embryo. Guided by these niches, we profiled cellular networks over space and time, not examinable using single-organ atlases. In so doing, we revealed tissue-specific fibroblast patterning from previously undescribed mesenchyme progenitors, early diversification of organ-specific blood capillaries and lymphatic vasculature, emergence of neural crest cell fates, the formation of placode- and neural crest-derived peripheral sensory neurons, and how tissue niches guide peripheral neuron maturation and axonal migration. The HDCA thus serves as a comprehensive step towards a comprehensive understanding of human prenatal development, and a template towards unravelling the biology of congenital disorders.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 19 matches between paragraphs and lines of code.

haniffalab/hdca-paper

License: none: the authors keep all their rights
State: the link is dead, verified on 28 September 2026
Evidence: found in the paper
Software Heritage: not archived
Found in: “Code”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 28 September 2026: the link is dead
  • 28 September 2026: the link is dead

theislab/idtrack

License: BSD-3-Clause
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: d8514ab51a494842f4278e769c9538f103592c7a, 20 January 2026
Languages: Python (72), Jupyter (38), Shell (8)
Size: 180 files, 118 scripts
Software Heritage: archived
Found in: “Code”
Holds: README, license file, environment (poetry.lock, pyproject.toml, docs/requirements.txt), tests, continuous integration, documentation, 38 notebooks
Not found: CITATION.cff
Tools: pandas (40 files), NumPy (31 files), Matplotlib (20 files), seaborn (18 files), anndata (7 files), NetworkX (5 files), SciPy (4 files), Scanpy (3 files), h5py (2 files), Biopython (1 file), PyTorch (1 file), Hugging Face Transformers (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
63 files

cellgeni/OMERO_tools

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: f4d9d3f1633265743936db2cfe48b53eda992272, 28 July 2026
Languages: Python (8)
Size: 14 files, 8 scripts
Software Heritage: not archived
Found in: the text, “Whole embryo: manual annotation of H&E images, a”
Holds: README, environment (environment.yml)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (6 files), pandas (6 files), Pillow (6 files), Scanpy (5 files), Squidpy (5 files), Matplotlib (4 files)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
9 files

haniffalab/cherita-react

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 740eaf25038452adf553d2528c0c8e8bcbd62906, 8 May 2026
Languages: JavaScript (27)
Size: 150 files, 27 scripts
Software Heritage: not archived
Found in: the text, “HDCA web portal integration with developmental d”
Holds: README, license file, CITATION.cff, tests, continuous integration, documentation
Not found: environment file
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
29 files

theislab/sctram

License: BSD-3-Clause
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: e570de5f9ec6dc783542b82c4f88b192677d4347, 10 April 2026
Languages: Python (125), Jupyter (53), Shell (6)
Size: 235 files, 184 scripts
Software Heritage: not archived
Found in: “Code”
Holds: README, license file, environment (poetry.lock, pyproject.toml, docs/requirements.txt), tests, continuous integration, documentation, 53 notebooks
Not found: CITATION.cff
Tools: anndata (18 files), Scanpy (17 files), NetworkX (16 files), NumPy (16 files), Matplotlib (15 files), pandas (15 files), seaborn (15 files), SciPy (5 files), PyTorch (4 files), scikit-learn (3 files), CuPy (1 file), scVelo (1 file), statsmodels (1 file), UMAP (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
35 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 5 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 129 scripts, each with its path and the digest of its content;
  • 19 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data and code availability

Data: Raw de novo whole embryo data will be available on EMBL-EBI ArrayExpress and BioImage Archive. All data accessions are currently private, and will be made available upon publication:

scRNAseq raw FASTQ and counts (embryo F137), E-MTAB-11504

scRNAseq raw FASTQ and counts (embryo F147), E-MTAB-11520

scRNAseq raw FASTQ and counts (embryo F158), E-MTAB-11911

Multiome raw FASTQ and counts (embryos F137, F147, F158), E-MTAB-16757

Spatial transcriptomics Visium raw FASTQs, counts, images (embryo SOB26), E-MTAB-14785

Spatial transcriptomics Xenium raw cell x gene counts post Xenium cell segmentation, images (embryos SOB25, SOB26), BioImage Archive, S-BIAD2878

The HDCA atlas will be accessible via an interactive web portal at https://cellatlas.io/studies/hdca/. This webportal is currently private and password protected, and will be made available upon publication. Through this portal, users will be able to visualise the dataset and download both raw and processed count matrices, as well as associated metadata and models.

Code: All code will be made publicly available at https://github.com/haniffalab/hdca-paper; the repository is currently private, and will be made available upon publication.

Additional repositories associated with this study include https://github.com/theislab/idtrack181 for biological identifier mapping and harmonization across heterogeneous data sources, and https://github.com/theislab/sctram45, which provides the trajectory-fidelity benchmarking framework introduced in this study.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, journal, dates, 94 authors, 5 funders, 189 references.

Cite

This paper

Webb, S., Rose, A., Xu, C., Steele, L., Kuri, M. A., Stephenson, E., Inecik, K., Jafree, D., Foster, A. R., Basurto-Lozada, D., Chipampe, N.-J., Pournara, A. V., Jacques, M.-A., To, K., Admane, C., Kritikaki, E., Chroscik, M. A., Horsfall, D., Foreman, J., . . . Haniffa, M. (2026). An integrated single cell and spatial omics atlas of human prenatal development. bioRxiv (preprint). https://doi.org/10.64898/2026.03.30.714220

BibTeX

@article{webb2026integrated,
author = {Webb, Simone and Rose, Antony and Xu, Chuan and Steele, Lloyd and Kuri, Masami A and Stephenson, Emily and Inecik, Kemal and Jafree, Daniyal and Foster, April R and Basurto-Lozada, Daniela and Chipampe, Nana-Jane and Pournara, Anna V and Jacques, Marc-Antoine and To, Ken and Admane, Chloe and Kritikaki, Efpraxia and Chroscik, Marta A and Horsfall, David and Foreman, Julia and Rademaker, Koen and Karjalainen, Julia and Laddach, Anna and Madad, Shaista and Lawrence, John EG and Kleshchevnikov, Vitalii and Lisgo, Steven and Lee, Jimmy Tsz Hang and Blevinal, Joséphine and Alqahtani, Ahlam and Makarchuk, Stanislaw and Jackson, Jake and Ucuncu, Ekin and Silva, Teresa P and Lorenzi, Valentina and Torabi, Fereshteh and Botting, Rachel A and Roberts, Kenny and Olabi, Bayanne and Chakala, Keerthi P and Dony, Leander and Dall’Aglio, Gabriele and Cujba, Ana-Maria and Whitfield, Holly J and Hale, Christine and Engelbert, Justin and Li, Tong and Prigmore, Elena and Gopee, Nusayhah Hudaa and Mather, Michael W and Talks, Ben and Fei, Lijiang and Pereira, Diana and Moldovan, Corina and Elmentaite, Rasa and Polanski, Krzysztof and Predeus, Alexander V and Mazin, Pasha and Rowe, Vicky and Bayraktar, Semih and Knight-Schrijver, Vincent R and Clatworthy, Menna and Kanemaru, Kazumasa and Cranley, James and Guo, Andrew and Sinha, Sanjay and McDonald, David and Filby, Andrew and Kerwin, Janet and He, Peng and Sequeira, Inês and Vento-Tormo, Roser and Luecken, Malte D. and Chédotal, Alain and Taylor, Deanne and Martin-Almedina, Silvia and Ostergaard, Pia and Behjati, Sam and Robertson, Elizabeth and Wilson, Nicola K and Marioni, John and Firth, Helen V and Wright, Caroline F and Pourquie, Olivier and Gambardella, Laure and Alexandre, Paula and Gottgens, Berthold and Meyer, Kerstin B and Vijayabaskar, MS and Lotfollahi, Mo and Teichmann, Sarah A and Theis, Fabian J and Progatzky, Fränze and Bayraktar, Omer Ali and Haniffa, Muzlifah},
title = {{An integrated single cell and spatial omics atlas of human prenatal development}},
journal = {bioRxiv (preprint)},
year = {2026},
month = apr,
publisher = {bioRxiv},
issn = {2692-8205},
doi = {10.64898/2026.03.30.714220},
url = {https://doi.org/10.64898/2026.03.30.714220}
}

RIS

TY - JOUR
AU - Webb, Simone
AU - Rose, Antony
AU - Xu, Chuan
AU - Steele, Lloyd
AU - Kuri, Masami A
AU - Stephenson, Emily
AU - Inecik, Kemal
AU - Jafree, Daniyal
AU - Foster, April R
AU - Basurto-Lozada, Daniela
AU - Chipampe, Nana-Jane
AU - Pournara, Anna V
AU - Jacques, Marc-Antoine
AU - To, Ken
AU - Admane, Chloe
AU - Kritikaki, Efpraxia
AU - Chroscik, Marta A
AU - Horsfall, David
AU - Foreman, Julia
AU - Rademaker, Koen
AU - Karjalainen, Julia
AU - Laddach, Anna
AU - Madad, Shaista
AU - Lawrence, John EG
AU - Kleshchevnikov, Vitalii
AU - Lisgo, Steven
AU - Lee, Jimmy Tsz Hang
AU - Blevinal, Joséphine
AU - Alqahtani, Ahlam
AU - Makarchuk, Stanislaw
AU - Jackson, Jake
AU - Ucuncu, Ekin
AU - Silva, Teresa P
AU - Lorenzi, Valentina
AU - Torabi, Fereshteh
AU - Botting, Rachel A
AU - Roberts, Kenny
AU - Olabi, Bayanne
AU - Chakala, Keerthi P
AU - Dony, Leander
AU - Dall’Aglio, Gabriele
AU - Cujba, Ana-Maria
AU - Whitfield, Holly J
AU - Hale, Christine
AU - Engelbert, Justin
AU - Li, Tong
AU - Prigmore, Elena
AU - Gopee, Nusayhah Hudaa
AU - Mather, Michael W
AU - Talks, Ben
AU - Fei, Lijiang
AU - Pereira, Diana
AU - Moldovan, Corina
AU - Elmentaite, Rasa
AU - Polanski, Krzysztof
AU - Predeus, Alexander V
AU - Mazin, Pasha
AU - Rowe, Vicky
AU - Bayraktar, Semih
AU - Knight-Schrijver, Vincent R
AU - Clatworthy, Menna
AU - Kanemaru, Kazumasa
AU - Cranley, James
AU - Guo, Andrew
AU - Sinha, Sanjay
AU - McDonald, David
AU - Filby, Andrew
AU - Kerwin, Janet
AU - He, Peng
AU - Sequeira, Inês
AU - Vento-Tormo, Roser
AU - Luecken, Malte D.
AU - Chédotal, Alain
AU - Taylor, Deanne
AU - Martin-Almedina, Silvia
AU - Ostergaard, Pia
AU - Behjati, Sam
AU - Robertson, Elizabeth
AU - Wilson, Nicola K
AU - Marioni, John
AU - Firth, Helen V
AU - Wright, Caroline F
AU - Pourquie, Olivier
AU - Gambardella, Laure
AU - Alexandre, Paula
AU - Gottgens, Berthold
AU - Meyer, Kerstin B
AU - Vijayabaskar, MS
AU - Lotfollahi, Mo
AU - Teichmann, Sarah A
AU - Theis, Fabian J
AU - Progatzky, Fränze
AU - Bayraktar, Omer Ali
AU - Haniffa, Muzlifah
TI - An integrated single cell and spatial omics atlas of human prenatal development
T2 - bioRxiv (preprint)
J2 - bioRxiv
PY - 2026
DA - 2026/04/01
SN - 2692-8205
PB - bioRxiv
DO - 10.64898/2026.03.30.714220
UR - https://doi.org/10.64898/2026.03.30.714220
ER -

CSL-JSON

{
"id": "10.64898/2026.03.30.714220",
"type": "article",
"title": "An integrated single cell and spatial omics atlas of human prenatal development",
"container-title": "bioRxiv (preprint)",
"author": [
{
"family": "Webb",
"given": "Simone"
},
{
"family": "Rose",
"given": "Antony"
},
{
"family": "Xu",
"given": "Chuan"
},
{
"family": "Steele",
"given": "Lloyd"
},
{
"family": "Kuri",
"given": "Masami A"
},
{
"family": "Stephenson",
"given": "Emily"
},
{
"family": "Inecik",
"given": "Kemal"
},
{
"family": "Jafree",
"given": "Daniyal"
},
{
"family": "Foster",
"given": "April R"
},
{
"family": "Basurto-Lozada",
"given": "Daniela"
},
{
"family": "Chipampe",
"given": "Nana-Jane"
},
{
"family": "Pournara",
"given": "Anna V"
},
{
"family": "Jacques",
"given": "Marc-Antoine"
},
{
"family": "To",
"given": "Ken"
},
{
"family": "Admane",
"given": "Chloe"
},
{
"family": "Kritikaki",
"given": "Efpraxia"
},
{
"family": "Chroscik",
"given": "Marta A"
},
{
"family": "Horsfall",
"given": "David"
},
{
"family": "Foreman",
"given": "Julia"
},
{
"family": "Rademaker",
"given": "Koen"
},
{
"family": "Karjalainen",
"given": "Julia"
},
{
"family": "Laddach",
"given": "Anna"
},
{
"family": "Madad",
"given": "Shaista"
},
{
"family": "Lawrence",
"given": "John EG"
},
{
"family": "Kleshchevnikov",
"given": "Vitalii"
},
{
"family": "Lisgo",
"given": "Steven"
},
{
"family": "Lee",
"given": "Jimmy Tsz Hang"
},
{
"family": "Blevinal",
"given": "Joséphine"
},
{
"family": "Alqahtani",
"given": "Ahlam"
},
{
"family": "Makarchuk",
"given": "Stanislaw"
},
{
"family": "Jackson",
"given": "Jake"
},
{
"family": "Ucuncu",
"given": "Ekin"
},
{
"family": "Silva",
"given": "Teresa P"
},
{
"family": "Lorenzi",
"given": "Valentina"
},
{
"family": "Torabi",
"given": "Fereshteh"
},
{
"family": "Botting",
"given": "Rachel A"
},
{
"family": "Roberts",
"given": "Kenny"
},
{
"family": "Olabi",
"given": "Bayanne"
},
{
"family": "Chakala",
"given": "Keerthi P"
},
{
"family": "Dony",
"given": "Leander"
},
{
"family": "Dall’Aglio",
"given": "Gabriele"
},
{
"family": "Cujba",
"given": "Ana-Maria"
},
{
"family": "Whitfield",
"given": "Holly J"
},
{
"family": "Hale",
"given": "Christine"
},
{
"family": "Engelbert",
"given": "Justin"
},
{
"family": "Li",
"given": "Tong"
},
{
"family": "Prigmore",
"given": "Elena"
},
{
"family": "Gopee",
"given": "Nusayhah Hudaa"
},
{
"family": "Mather",
"given": "Michael W"
},
{
"family": "Talks",
"given": "Ben"
},
{
"family": "Fei",
"given": "Lijiang"
},
{
"family": "Pereira",
"given": "Diana"
},
{
"family": "Moldovan",
"given": "Corina"
},
{
"family": "Elmentaite",
"given": "Rasa"
},
{
"family": "Polanski",
"given": "Krzysztof"
},
{
"family": "Predeus",
"given": "Alexander V"
},
{
"family": "Mazin",
"given": "Pasha"
},
{
"family": "Rowe",
"given": "Vicky"
},
{
"family": "Bayraktar",
"given": "Semih"
},
{
"family": "Knight-Schrijver",
"given": "Vincent R"
},
{
"family": "Clatworthy",
"given": "Menna"
},
{
"family": "Kanemaru",
"given": "Kazumasa"
},
{
"family": "Cranley",
"given": "James"
},
{
"family": "Guo",
"given": "Andrew"
},
{
"family": "Sinha",
"given": "Sanjay"
},
{
"family": "McDonald",
"given": "David"
},
{
"family": "Filby",
"given": "Andrew"
},
{
"family": "Kerwin",
"given": "Janet"
},
{
"family": "He",
"given": "Peng"
},
{
"family": "Sequeira",
"given": "Inês"
},
{
"family": "Vento-Tormo",
"given": "Roser"
},
{
"family": "Luecken",
"given": "Malte D."
},
{
"family": "Chédotal",
"given": "Alain"
},
{
"family": "Taylor",
"given": "Deanne"
},
{
"family": "Martin-Almedina",
"given": "Silvia"
},
{
"family": "Ostergaard",
"given": "Pia"
},
{
"family": "Behjati",
"given": "Sam"
},
{
"family": "Robertson",
"given": "Elizabeth"
},
{
"family": "Wilson",
"given": "Nicola K"
},
{
"family": "Marioni",
"given": "John"
},
{
"family": "Firth",
"given": "Helen V"
},
{
"family": "Wright",
"given": "Caroline F"
},
{
"family": "Pourquie",
"given": "Olivier"
},
{
"family": "Gambardella",
"given": "Laure"
},
{
"family": "Alexandre",
"given": "Paula"
},
{
"family": "Gottgens",
"given": "Berthold"
},
{
"family": "Meyer",
"given": "Kerstin B"
},
{
"family": "Vijayabaskar",
"given": "MS"
},
{
"family": "Lotfollahi",
"given": "Mo"
},
{
"family": "Teichmann",
"given": "Sarah A"
},
{
"family": "Theis",
"given": "Fabian J"
},
{
"family": "Progatzky",
"given": "Fränze"
},
{
"family": "Bayraktar",
"given": "Omer Ali"
},
{
"family": "Haniffa",
"given": "Muzlifah"
}
],
"container-title-short": "bioRxiv",
"DOI": "10.64898/2026.03.30.714220",
"ISSN": "2692-8205",
"publisher": "bioRxiv",
"URL": "https://doi.org/10.64898/2026.03.30.714220",
"issued": {
"date-parts": [
[
2026,
4,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: Squidpy, CuPy, UMAP, 12 other tools, 3 references
[2] doi:10.1038/s41467-026-68596-w [code]
Spatial cartography of human thymus enables the geopositioning of lineage transcription factors in rare mimetic thymic epithelial cells.
Journal: Nature communications
In common: Squidpy, anndata, Scanpy, 9 other tools, 5 references
[3] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: CuPy, Biopython, UMAP, 12 other tools, 2 references
[4] doi:10.1016/j.celrep.2026.117270 [code]
Blocking apoptosis promotes survival and alters developmental dynamics of human retinal ganglion cells in retinal organoids.
Journal: Cell reports
In common: UMAP, anndata, Scanpy, 7 other tools, 6 references
[5] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: scVelo, UMAP, anndata, 10 other tools, 3 references
[6] doi:10.1016/j.crmeth.2026.101342 [code]
Interpretable learning of temporal cellular dynamics from single-cell data.
Journal: Cell reports methods
In common: scVelo, UMAP, anndata, 11 other tools, 1 reference
[7] doi:10.21203/rs.3.rs-9676637/v1 [code]
A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
Journal: Research Square (preprint)
In common: Squidpy, UMAP, anndata, 10 other tools, 2 references
[8] doi:10.1038/s44320-026-00208-7 [code]
Interpretable deep generative ensemble learning for single-cell omics with Hydra.
Journal: Molecular systems biology
In common: UMAP, anndata, Scanpy, 8 other tools, 5 references
[9] doi:10.1016/j.xgen.2026.101217 [code]
ProtoCloud: A prototypical self-explaining model for single-cell analysis.
Journal: Cell genomics
In common: UMAP, anndata, Scanpy, 8 other tools, 5 references
[10] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: Squidpy, UMAP, anndata, 11 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.