An integrated single cell and spatial omics atlas of human prenatal development
The 19 matches
- [1] § Methods › Whole embryo: scRNAseq integration, batch correction and visualisation ↔ reproducibility/modules/module_generate/submodule_synthetic_simple/example_operations.ipynb, lines 292–373 · score 0.84 · n_neighbors, sc.pp.neighbors, sc.tl.umap, Scanpy, dimensionality, heatmaps
- [2] § Methods › Community data: Realignment and gene space harmonisation ↔ idtrack/_harmonize_features.py, lines 22–82 · score 0.81 · failed conversions, Gene symbols, gene identifiers, IDTrack, conflicts, audit
- [3] § Methods › Whole embryo: scRNAseq integration, batch correction and visualisation ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 1793–1810 · score 0.81 · n_neighbors, sc.pp.neighbors, sc.tl.umap, Scanpy, scVI, atlas
- [4] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › scTRAM: Trajectory-specific fidelity evaluation ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/incremental_analysis.ipynb, lines 70–155 · score 0.81 · graph edit distance, spectral distance, trustworthiness, Spearman, concordance, stress
- [5] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › Integrated benchmarking framework ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/helper/suo_scib_iterative.py, lines 25–133 · score 0.79 · min max scaled, scIB, aggregate score, batch correction, conservation, trajectory
- [6] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 198–252 · score 0.78 · TarDis, scPoli, scTRAM, scVI, Harmony, PCA
- [7] § Methods › Community data: Integration with scVI ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 198–252 · score 0.76 · TarDis, scPoli, scTRAM, scVI, Harmony, PCA
- [8] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › Integrated benchmarking framework ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 254–290 · score 0.75 · scTRAM, scIB, aggregate score, batch correction, topology, conservation
- [9] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › scIB: Conventional integration quality assessment ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/incremental_analysis.ipynb, lines 241–324 · score 0.70 · graph connectivity, scIB, cLISI, iLISI, isolated, silhouette
- [10] § Methods › Community data: Benchmarking integration and trajectories with scIB and scTRAM › scTRAM: Trajectory-specific fidelity evaluation ↔ reproducibility/modules/module_evaluate/example_operations_adjacency.ipynb, lines 92–119 · score 0.68 · graph edit distance, spectral distance, adjacency, correlations, scores, metric
- [11] § Methods › HDCA: Integration and annotation harmonisation of the final integrated community and whole embryo data ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/helper/model_training.py, lines 15–97 · score 0.66 · gene likelihood, batch_key, trained, covariate, latent, layer
- [12] § Results › Benchmarking atlas integration with trajectory fidelity ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 380–453 · score 0.66 · TarDis, scPoli, scVI, Harmony, PCA, scTRAM
- [13] § Methods › scRNAseq gene expression and UMAP visualisation ↔ reproducibility/modules/module_generate/submodule_synthetic_simple/example_operations.ipynb, lines 292–373 · score 0.64 · sc.pl.umap, Scanpy, pp, heatmaps, matrices
- [14] § Methods › HDCA: Integration and annotation harmonisation of the final integrated community and whole embryo data ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/incremental_analysis.ipynb, lines 70–155 · score 0.63 · Normalised Mutual Information, macrophage, myeloid, concordance, Suo, neighbour
- [15] § Results › Benchmarking atlas integration with trajectory fidelity ↔ reproducibility/local/experiments/sctram_figures_only.ipynb, lines 1793–1810 · score 0.58 · ground truth, UMAP, HSC, MPP, scPoli, scTRAM
- [16] § Methods › Community data: Highly variable genes selection pipeline ↔ idtrack/_harmonize_features.py, lines 22–82 · score 0.53 · harmonised genes, IDTrack, single cell, logs, conversion, downstream
- [17] § Methods › Community data: Public human developmental single cell data acquisition and preprocessing ↔ idtrack/_harmonize_features.py, lines 876–941 · score 0.52 · trivial, age, provenance, donor, concatenated, schema
- [18] § Methods › Community data: Integration with scVI ↔ idtrack/_harmonize_features.py, lines 943–979 · score 0.51 · AnnData, Scanpy, uns, post, validation, genes
- [19] § Methods › Whole embryo: Xenium preprocessing and analysis ↔ reproducibility/server_sync/experiments/_deprecated_suo_incremental/helper/model_training.py, lines 15–97 · score 0.50 · categorical covariate, unlabelled, SCANVI, latent, layers, Batch
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 1,813 lines · 57 KB · BSD-3-Clause · 6 matches
- # %% [markdown]
- # # Imports
- # %%
- 0
- # %%
- import sys
- import os
- import gc
- import networkx as nx
- import numpy as np
- import pandas as pd
- import tqdm
- # %%
- import anndata as ad
- import scanpy as sc
- # import scvelo as scv
- # import cellrank as cr
- # import scanpy.external as sce
- # scv.settings.verbosity = 3
- # cr.settings.verbosity = 2
- # sc.settings.verbose = 3
- # %%
- working_directory = "/Users/kemalinecik/git_nosync/sctram"
- sys.path.append(working_directory)
- sys.path.append("/Users/kemalinecik/git_nosync/sctram/reproducibility/server_sync/experiments")
- from incremental_training.helper.constants import metrics_scib, metric_direction_dict
- from development_atlas_figure.helper.get_hdca_trainings_with_paths import get_hdca_trainings_with_paths
- from development_atlas_figure.helper.pickle_operations import save_pickle_bundle, load_pickle_bundle
- from development_atlas_figure.helper.csv_operations import (
- harmonize_union_with_report,
- analyze_column_consistency,
- combine_dataframes_on_columns
- )
- from sctram.api._lower_level import TrajectoryEvaluationAPI
- from sctram.generate.real import sc_suo_developmental_complete
- from sctram.input import InputTrajectories
- # %%
- %matplotlib inline
- %config InlineBackend.figure_format='retina'
- import pickle
- import matplotlib.pyplot as plt
- import seaborn as sns
- import colorcet as cc
- import matplotlib.patheffects as path_effects
- from adjustText import adjust_text
- from matplotlib import gridspec
- from networkx.drawing.nx_agraph import graphviz_layout
- import matplotlib as mpl
- import matplotlib.patches as mpatches
- from matplotlib.ticker import FormatStrFormatter
- from adjustText import adjust_text
- import matplotlib.patheffects as path_effects
- import matplotlib.cm as cm
- from plottable import ColumnDefinition, Table
- from plottable.cmap import normed_cmap
- from matplotlib.ticker import MaxNLocator
- from plottable.plots import bar
- _rcparams_path = os.path.join(working_directory, "reproducibility/figure_rcparams/rcparams.pickle")
- with open(_rcparams_path, "rb") as file:
- _rcparams = pickle.load(file)
- plt.rcParams.update(_rcparams)
- plt.rcParams["font.family"] = "Arial"
- plt.rcParams["font.sans-serif"] = ["Arial"]
- # %%
- temp_move = "/Users/kemalinecik/Downloads/temp_move"
- # %% [markdown]
- # how many
- # %%
- (
- df_hdca_component_adata_paths,
- df_hdca_lvl3_adata_paths,
- out_df,
- ground_truth_trajectories,
- ground_truth_trajectories_annotation_level,
- trajectory_subset_dict
- ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
- ground_truth_trajectories.keys()
- # %%
- InputTrajectories(ground_truth_trajectories).get_trajectory("germline", False)
- # %%
- InputTrajectories(ground_truth_trajectories).get_trajectory("Haem_ref", False)
- # %%
- InputTrajectories(ground_truth_trajectories).get_trajectory("pns_neuro", False)
- # %%
- fig, ax = InputTrajectories(ground_truth_trajectories).get_trajectory("germline", False).plot_trajectory(
- title="",
- figsize=(12, 1.65),
- label_offset=0,
- return_ax=True
- )
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/graph_germline.pdf")
- # %%
- fig, ax = InputTrajectories(ground_truth_trajectories).get_trajectory("Haem_ref", False).plot_trajectory(
- title="",
- figsize=(12, 7),
- label_offset=0,
- return_ax=True
- )
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/graph_Haem_ref.pdf")
- # %%
- fig, ax = InputTrajectories(ground_truth_trajectories).get_trajectory("pns_neuro", False).plot_trajectory(
- title="",
- figsize=(12, 2),
- label_offset=0,
- return_ax=True
- )
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/graph_pns_neuro.pdf")
- # %% [markdown]
- # # Figure
- # %%
- def logistic_invert_columns(df, col=None, k=1.0):
- """
- Apply logistic inversion scaling to a column or entire DataFrame.
- Parameters
- ----------
- df : pd.DataFrame or pd.Series
- Input DataFrame or Series.
- col : str or None, default=None
- Column name to transform. If None, apply to all numeric columns.
- k : float, default=1.0
- Steepness of the sigmoid. Larger k -> sharper curve.
- Returns
- -------
- pd.DataFrame or pd.Series
- Transformed DataFrame/Series with values in (0, 1).
- """
- if isinstance(df, pd.Series): # if a single column (Series)
- return 1 / (1 + np.exp(k * (df - df.mean())))
- if col is not None: # single column of DataFrame
- df = df.copy()
- x = df[col]
- df[col] = 1 / (1 + np.exp(k * (x - x.mean())))
- return df
- # all numeric columns
- df = df.copy()
- for c in df.select_dtypes(include=np.number).columns:
- x = df[c]
- df[c] = 1 / (1 + np.exp(k * (x - x.mean())))
- return df
- def logistic_invert_rowwise(df, k=1.0):
- """
- Apply logistic inversion scaling row-wise to all numeric columns.
- Parameters
- ----------
- df : pd.DataFrame
- Input DataFrame.
- k : float, default=1.0
- Steepness of the sigmoid. Larger k -> sharper curve.
- Returns
- -------
- pd.DataFrame
- Transformed DataFrame with values in (0, 1).
- """
- df = df.copy()
- numeric_cols = df.select_dtypes(include=np.number).columns
- # Apply rowwise
- def transform_row(row):
- mean_val = row[numeric_cols].mean()
- return 1 / (1 + np.exp(k * (row[numeric_cols] - mean_val)))
- df[numeric_cols] = df[numeric_cols].apply(transform_row, axis=1)
- return df
- # %%
- def rank_metrics(df, metric_direction_dict, metric_index_pos, method='min'):
- ranked_df = pd.DataFrame(index=df.index, columns=df.columns, dtype=float)
- for indices, row in df.iterrows():
- metric = indices[metric_index_pos]
- direction = metric_direction_dict.get(metric)
- if direction not in ('increasing', 'decreasing'):
- raise ValueError
- numeric_row = pd.to_numeric(row, errors='coerce')
- if numeric_row.isna().any():
- ranked_df.loc[indices] = np.nan
- continue
- if direction == 'increasing':
- ranks = numeric_row.rank(method='dense', ascending=False)
- else:
- ranks = numeric_row.rank(method='dense', ascending=True)
- # ranks = 1 - pd.Series(MinMaxScaler().fit_transform(row.values.reshape(-1, 1)).flatten(), index=row.index)
- ranked_df.loc[indices] = ranks.astype('Float64')
- return ranked_df
- result_df = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_result_df_dataframe.pkl"))
- df_pivot = result_df.pivot(index=["trajectory_class", 'trajectory', 'path', 'metric'], columns='representation', values='score')
- df_pivot_rank = rank_metrics(df_pivot, metric_direction_dict, metric_index_pos=3, method='min')
- df_pivot_rank_mean = df_pivot_rank.groupby(level=1).mean()
- rename_model_name_dict = {
- 'pca': 'PCA',
- 'X_pca': 'PCA',
- 'harmony': 'Harmony',
- 'invae': 'inVAE',
- 'scanvi': 'scANVI',
- 'scvi': 'scVI',
- 'scpoli': 'scPoli',
- 'scanoroma': 'Scanoroma',
- 'tardis_1': 'DELETE ME',
- 'tardis_2': 'TarDis',
- 'tardis': 'TarDis',
- }
- col_order = ['Scanoroma', 'Harmony', 'TarDis', 'PCA', 'scPoli', 'scVI']
- sctram_summary = (logistic_invert_rowwise(df_pivot_rank).groupby(level=[2, 3]).mean().groupby(level=0).mean())
- _sctram_summary = (pd.DataFrame(logistic_invert_rowwise(df_pivot_rank).groupby(level=3).mean().mean(), columns=["Total scTRAM"]).T)
- sctram_summary = pd.concat([sctram_summary, _sctram_summary], axis=0)
- sctram_summary = sctram_summary.rename(
- index={"adjacency": "Adjacency-based", "embedding": "Embedding-based", "pseudotime": "Pseudotime-based"},
- columns=rename_model_name_dict,
- )[col_order]
- sctram_summary.index.name = None
- sctram_summary.columns.name = "Representation"
- sctram_summary
- # %%
- # sctram_summary = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_sctram_summary_dataframe.pkl")) # invert back to rank again
- scib_summary = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_scib_summary_dataframe.pkl"))
- replacements = {
- "Pseudotime-based": "Cell Order\ncontinuity",
- "Embedding-based": "Trajectory\ndominance",
- "Adjacency-based": "Topology\nconservation",
- "Bio conservation": "Biological\nconservation",
- "Batch correction": "Batch\ncorrection",
- }
- summary = pd.concat([scib_summary, sctram_summary], keys=["scIB", "scTRAM"], axis=0)
- summary.index.names = ["Source", "Metric Group"]
- summary.columns.name = "Representation"
- new_index_tuples = [
- ("Aggregate score", metric.split("Total ")[1]) if metric.startswith("Total ") else (src, metric) for src, metric in summary.index
- ]
- summary.index = pd.MultiIndex.from_tuples(new_index_tuples, names=summary.index.names)
- weights = { # just provisional for now
- "scTRAM": 0.25,
- "scIB": 0.75,
- }
- row_ts = summary.loc[("Aggregate score", "scTRAM")]
- row_si = summary.loc[("Aggregate score", "scIB")]
- combined = weights["scTRAM"] * row_ts + weights["scIB"] * row_si
- combined_df = pd.DataFrame(combined).T
- combined_df.index = pd.MultiIndex.from_tuples([("Aggregate score", "Total")], names=summary.index.names)
- summary = pd.concat([summary, combined_df])
- summary = summary.sort_index(ascending=False)
- summary = summary.sort_values(by=[("Aggregate score", "Total")], axis=1, ascending=False).astype(np.float64)
- summary.rename(index=replacements, level="Metric Group", inplace=True)
- summary_T = summary.T
- summary
- # %%
- df_plot = summary_T.reset_index().copy()
- df_plot_dict = {f"{lvl0}\n{lvl1}": (lvl0, lvl1) for (lvl0, lvl1) in df_plot.columns}
- df_plot.columns = [f"{lvl0}\n{lvl1}" for (lvl0, lvl1) in df_plot.columns]
- cmap_fn = lambda col_data: normed_cmap(col_data, cmap=cm.PRGn, num_stds=2.5)
- col_defs = []
- for col in df_plot.columns:
- # Create a “normalized color map” for this column:
- lvl0, lvl1 = df_plot_dict[col]
- if lvl0 == "Aggregate score" or lvl0 == "Representation":
- continue
- cmap = cmap_fn(df_plot[col])
- cd = ColumnDefinition(
- name=col, # Must exactly match the flattened column name in df_plot
- width=0.4, # width of each column in inches (tweak as needed)
- cmap=cmap, # cell background colormap (normed by this column’s values)
- text_cmap=cmap, # text color also follows same colormap
- textprops={
- "ha": "center",
- "bbox": {"boxstyle": "circle", "pad": 0},
- },
- group=lvl0,
- title=lvl1,
- formatter="{:.3f}".format, # three decimal places
- )
- col_defs.append(cd)
- i = 0
- for col in df_plot.columns:
- # Create a “normalized color map” for this column:
- lvl0, lvl1 = df_plot_dict[col]
- if lvl0 != "Aggregate score" or lvl0 == "Representation":
- continue
- cmap = cmap_fn(df_plot[col])
- i += 1
- c_min = min(df_plot[col])
- c_max = max(df_plot[col])
- cmap = cmap_fn(df_plot[col])
- cd = ColumnDefinition(
- name=col,
- width=0.3,
- plot_fn=bar,
- plot_kw={
- "cmap": cmap, #cm.YlGnBu,
- "plot_bg_bar": False,
- # "xlim": [c_min, c_max],
- "annotate": True,
- "height": 0.7,
- "formatter": "{:.3f}",
- "textprops": {"fontsize": 10}
- },
- group=lvl0,
- title=lvl1,
- border="left" if i == 0 else None,
- )
- col_defs.append(cd)
- col_defs += [ColumnDefinition(name="Representation\n", title="", width=0.3, textprops={"ha": "left", "weight": "bold"})]
- # Allow to manipulate text post-hoc (in illustrator)
- with mpl.rc_context({"svg.fonttype": "none"}):
- fig, ax = plt.subplots(figsize=(len(summary_T.columns) * 1.3, 3 + 0.25 * len(summary_T.index)))
- tab = Table(
- df_plot,
- cell_kw={
- "linewidth": 0,
- "edgecolor": "k",
- },
- row_dividers=True,
- footer_divider=True,
- textprops={"fontsize": 10, "ha": "center"},
- row_divider_kw={"linewidth": 0.25, "linestyle": "-"},
- col_label_divider_kw={"linewidth": 1, "linestyle": "-"},
- column_border_kw={"linewidth": 1, "linestyle": "-"},
- column_definitions=col_defs,
- even_row_color="#97ffff",
- odd_row_color="#ffffff",
- ax=ax,
- index_col="Representation\n",
- ).autoset_fontcolors(colnames=df_plot.columns)
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/table.pdf")
- # %% [markdown]
- # # Figure
- # %%
- result_df = pd.read_pickle(os.path.join(temp_move, "_experiment_development_atlas_result_df_dataframe.pkl"))
- df_pivot = result_df.pivot(index=["trajectory_class", 'trajectory', 'path', 'metric'], columns='representation', values='score')
- df_pivot_rank = rank_metrics(df_pivot, metric_direction_dict, metric_index_pos=3, method='min')
- df_pivot_rank_mean = logistic_invert_rowwise(df_pivot_rank).groupby(level=1).mean()
- rename_model_name_dict = {
- 'pca': 'PCA',
- 'X_pca': 'PCA',
- 'harmony': 'Harmony',
- 'invae': 'inVAE',
- 'scanvi': 'scANVI',
- 'scvi': 'scVI',
- 'scpoli': 'scPoli',
- 'scanoroma': 'Scanoroma',
- 'tardis_1': 'DELETE ME',
- 'tardis_2': 'TarDis',
- 'tardis': 'TarDis',
- }
- col_order = ['Scanoroma', 'Harmony', 'TarDis', 'PCA', 'scPoli', 'scVI']
- def _(df):
- df = df.rename(columns=rename_model_name_dict)
- df = df[col_order]
- fig_height=4
- n_cols = df.shape[1]
- col_width = 0.7 # Width per column in inches
- fig_width = n_cols * col_width
- # fig_height = 8 # Fixed height for better vertical readability
- # Create figure with dynamic width
- plt.figure(figsize=(fig_width, fig_height))
- # Create boxplot with refined styling
- bp = sns.boxplot(data=df,
- width=0.8,
- showfliers=False,
- boxprops=dict(facecolor='lightgray', edgecolor='black', linewidth=0.81),
- whiskerprops=dict(color='black', linewidth=0.81),
- medianprops=dict(color='darkred', linewidth=0.81),
- capprops=dict(color='black', linewidth=0.81),
- saturation=1,
- zorder=1)
- # Add swarmplot with sophisticated aesthetics
- sp = sns.swarmplot(data=df,
- size=4,
- palette='dark:black',
- edgecolor='white',
- linewidth=0.6,
- alpha=0.8,
- marker='o',
- zorder=2)
- # Enhance axis styling
- plt.xticks(rotation=90, ha='center', fontsize=10)
- plt.yticks(fontsize=10)
- plt.xlabel('Representation', fontsize=10, labelpad=10)
- plt.ylabel("Performance", fontsize=10, labelpad=10)
- # plt.gca().set_ylim(1.75, 5.5)
- plt.gca().yaxis.set_major_formatter(FormatStrFormatter('%.2f'))
- # Add light grid and adjust borders
- bp.grid(axis='y', linestyle='--', alpha=0.4)
- sns.despine(left=True, trim=False)
- # Adjust layout with professional spacing
- plt.tight_layout(rect=[0.05, 0.05, 0.95, 0.95])
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/modelcomp_boxplot.pdf")
- _(df_pivot_rank_mean)
- # %%
- df_pivot_rank_mean = logistic_invert_rowwise(df_pivot_rank).groupby(level=[0, 1]).mean()
- df_pivot_rank_mean = df_pivot_rank_mean.rename(columns=rename_model_name_dict)
- df_pivot_rank_mean = df_pivot_rank_mean[col_order]
- df_long = df_pivot_rank_mean.reset_index().melt(id_vars=['trajectory_class', 'trajectory'], var_name='representation', value_name='score')
- # %%
- # match boxplot’s height & per-column width
- fig_height = 3
- n_cols = df_long["representation"].nunique()
- col_width = 1.4
- fig_width = n_cols * col_width
- n_tc = len(np.unique(df_long["trajectory_class"]))
- fig, ax = plt.subplots(figsize=(fig_width, fig_height))
- # draw bars
- bp = sns.barplot(
- data=df_long,
- x="representation",
- y="score",
- hue="trajectory_class",
- palette=sns.color_palette("Greys", n_colors=n_tc+2),
- width=0.8,
- saturation=1,
- errorbar="sd",
- capsize=0.05,
- err_kws={'color': 'black', 'linewidth': 0.81},
- edgecolor="black",
- linewidth=0.81,
- zorder=1,
- ax=ax
- )
- # enforce edge styling
- for patch in bp.patches:
- patch.set_edgecolor("black")
- patch.set_linewidth(0.81)
- # axes
- ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
- ax.tick_params(axis="y", labelsize=10)
- ax.set_xlabel('Representation', fontsize=10, labelpad=10)
- plt.ylabel("Performance", fontsize=10, labelpad=10)
- ax.set_ylim(0.1, None)
- ax.yaxis.set_major_formatter(FormatStrFormatter('%.2f'))
- # grid + despine
- ax.grid(axis='y', linestyle='--', alpha=0.4)
- sns.despine(left=True, trim=False)
- # remove the old legend
- ax.legend_.remove()
- # build new, professional legend
- handles, labels = ax.get_legend_handles_labels()
- slim_handles = [
- mpatches.Patch(
- facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
- edgecolor="black",
- linewidth=0.5 # thinner border
- )
- for h in handles
- ]
- leg = ax.legend(
- slim_handles,
- labels,
- title="Trajectory Group",
- loc="center left",
- bbox_to_anchor=(1.02, 0.5),
- frameon=False,
- alignment="left",
- title_fontsize=8,
- fontsize=8,
- borderaxespad=0.5,
- handletextpad=0.4,
- labelspacing=0.5,
- )
- # tighten layout (accounts for legend)
- plt.tight_layout(rect=[0, 0, 0.85, 1])
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/modelcomp_barplot1.pdf")
- # %%
- df_pivot_rank_mean = logistic_invert_rowwise(df_pivot_rank).groupby(level=[2, 1]).mean()
- df_pivot_rank_mean = df_pivot_rank_mean.rename(columns=rename_model_name_dict)
- df_pivot_rank_mean = df_pivot_rank_mean[col_order]
- df_long = df_pivot_rank_mean.reset_index().melt(id_vars=['path', 'trajectory'], var_name='representation', value_name='score')
- # %%
- label_mapping = {
- "adjacency": 'Topology conservation',
- "embedding": 'Trajectory dominance',
- "pseudotime": 'Cell Order continuity',
- }
- # match boxplot’s height & per-column width
- fig_height = 3
- n_cols = df_long["representation"].nunique()
- col_width = 1.4
- fig_width = n_cols * col_width
- fig, ax = plt.subplots(figsize=(fig_width, fig_height))
- # draw bars
- bp = sns.barplot(
- data=df_long,
- x="representation",
- y="score",
- hue="path",
- palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
- width=0.8,
- saturation=1,
- errorbar="sd",
- capsize=0.05,
- err_kws={'color': 'black', 'linewidth': 0.81},
- edgecolor="black",
- linewidth=0.81,
- zorder=1,
- ax=ax
- )
- # enforce edge styling
- for patch in bp.patches:
- patch.set_edgecolor("black")
- patch.set_linewidth(0.81)
- # axes
- ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
- ax.tick_params(axis="y", labelsize=10)
- ax.set_xlabel('Representation', fontsize=10, labelpad=10)
- plt.ylabel("Performance", fontsize=10, labelpad=10)
- ax.set_ylim(0.1, None)
- ax.yaxis.set_major_formatter(FormatStrFormatter('%.2f'))
- # grid + despine
- ax.grid(axis='y', linestyle='--', alpha=0.4)
- sns.despine(left=True, trim=False)
- # remove the old legend
- ax.legend_.remove()
- # build new, professional legend
- handles, labels = ax.get_legend_handles_labels()
- new_labels = [label_mapping.get(l, l) for l in labels]
- slim_handles = [
- mpatches.Patch(
- facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
- edgecolor="black",
- linewidth=0.5 # thinner border
- )
- for h in handles
- ]
- leg = ax.legend(
- slim_handles,
- new_labels,
- title="Metric Group",
- loc="center left",
- bbox_to_anchor=(1.02, 0.5),
- alignment="left",
- frameon=False,
- title_fontsize=8,
- fontsize=8,
- borderaxespad=0.5,
- handletextpad=0.4,
- labelspacing=0.5,
- )
- # tighten layout (accounts for legend)
- plt.tight_layout(rect=[0, 0, 0.85, 1])
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/modelcomp_barplot2.pdf")
- # %% [markdown]
- # # Figure 2
- # %%
- result_df_main = pd.read_pickle(os.path.join(temp_move, f"df_boots_fig2_calculations.pkl"))
- (
- df_hdca_component_adata_paths,
- df_hdca_lvl3_adata_paths,
- out_df,
- ground_truth_trajectories,
- ground_truth_trajectories_annotation_level,
- trajectory_subset_dict
- ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
- # %%
- from math import sqrt
- def rank_metrics(df, metric_direction_dict, metric_index_pos, method='min'):
- ranked_df = pd.DataFrame(index=df.index, columns=df.columns, dtype=float)
- for indices, row in df.iterrows():
- metric = indices[metric_index_pos]
- direction = metric_direction_dict.get(metric)
- if direction not in ('increasing', 'decreasing'):
- raise ValueError
- numeric_row = pd.to_numeric(row, errors='coerce')
- if numeric_row.isna().any():
- ranked_df.loc[indices] = np.nan
- continue
- if direction == 'increasing':
- ranks = numeric_row.rank(method='dense', ascending=False)
- else:
- ranks = numeric_row.rank(method='dense', ascending=True)
- # ranks = 1 - pd.Series(MinMaxScaler().fit_transform(row.values.reshape(-1, 1)).flatten(), index=row.index)
- ranked_df.loc[indices] = ranks.astype('Float64')
- return ranked_df
- def plot_trajectory(
- G,
- title="trajectory_name",
- figsize=None,
- font_size=10,
- node_size=300,
- node_color="#8DA0CB", # Modern muted blue
- cmap="viridis",
- color_nodes_by_position=False,
- # Edge styling parameters
- edge_width=1,
- edge_color="gray",
- color_edges_by_weight=False,
- edge_weight_attribute="weight",
- edge_cmap="plasma",
- edge_vmin=None,
- edge_vmax=None,
- edge_colorbar=True,
- edge_labels=False,
- edge_label_format=".2f",
- edge_label_font_size=8,
- edge_label_color="black",
- edge_label_offset=None,
- # Layout parameters
- label_offset=None,
- title_y=1.02,
- layout_args="-Grankdir=LR -Gnodesep=0.25 -Granksep=0.25",
- file_name="blabla"
- ):
- pos = nx.nx_agraph.graphviz_layout(G, prog="dot", args=layout_args)
- x_coords = [pos[node][0] for node in G.nodes()] if G.nodes() else []
- y_coords = [pos[node][1] for node in G.nodes()] if G.nodes() else []
- # Create actual figure with determined size
- fig = plt.figure(figsize=figsize)
- ax = fig.add_subplot(111)
- # Node coloring logic
- if color_nodes_by_position:
- try:
- min_x, max_x = min(x_coords), max(x_coords)
- color_vals = [(x - min_x) / (max_x - min_x) if max_x != min_x else 0.5 for x in x_coords]
- node_color = plt.get_cmap(cmap)(color_vals)
- except Exception as e:
- print(f"Warning: Positional node coloring failed - {e}. Using base color.")
- # Edge coloring logic
- if color_edges_by_weight:
- weights = [G[u][v].get(edge_weight_attribute, 1.0) for u, v in G.edges]
- # Separate NaN and valid weights
- valid_weights = [w for w in weights if not np.isnan(w)]
- if valid_weights:
- edge_vmin = edge_vmin if edge_vmin is not None else min(valid_weights)
- edge_vmax = edge_vmax if edge_vmax is not None else max(valid_weights)
- norm = mpl.colors.Normalize(vmin=edge_vmin, vmax=edge_vmax)
- cmap_edge = plt.get_cmap(edge_cmap)
- # Create edge colors, using gray with alpha=0.712 for NaN weights
- edge_colors = []
- edge_alphas = []
- for w in weights:
- if np.isnan(w):
- edge_colors.append("lightgray")
- edge_alphas.append(0.712)
- else:
- edge_colors.append(cmap_edge(norm(w)))
- edge_alphas.append(1.0)
- sm = plt.cm.ScalarMappable(norm=norm, cmap=cmap_edge)
- sm.set_array(valid_weights)
- else:
- edge_colors = "lightgray"
- edge_alphas = 0.712
- sm = None
- else:
- edge_colors = edge_color
- edge_alphas = 1.0
- sm = None
- # Label positioning offsets
- if label_offset is None:
- label_offset = 0
- if edge_label_offset is None:
- edge_label_offset = 0
- # Calculate final label positions
- label_pos = {n: (x, y + label_offset) for n, (x, y) in pos.items()}
- # Draw nodes and edges
- nx.draw_networkx_nodes(
- G, pos, ax=ax, node_size=node_size, node_color=node_color, edgecolors="black", linewidths=0.8, alpha=1.0
- )
- node_radius = sqrt(node_size) / 2.0
- if color_edges_by_weight and isinstance(edge_alphas, list):
- # Draw edges individually when we have different alpha values
- for i, (u, v) in enumerate(G.edges()):
- nx.draw_networkx_edges(
- G,
- pos,
- edgelist=[(u, v)],
- ax=ax,
- edge_color=[edge_colors[i]],
- width=edge_width,
- alpha=edge_alphas[i],
- connectionstyle="arc3",
- arrows=False,
- )
- else:
- nx.draw_networkx_edges(
- G,
- pos,
- ax=ax,
- edge_color=edge_colors,
- width=edge_width,
- alpha=edge_alphas if not isinstance(edge_alphas, list) else 1.0,
- connectionstyle="arc3",
- arrows=False,
- )
- # Edge label placement
- if edge_labels:
- for u, v in G.edges():
- if edge_weight_attribute not in G[u][v]:
- continue
- weight = G[u][v][edge_weight_attribute]
- x1, y1 = pos[u]
- x2, y2 = pos[v]
- dx, dy = x2 - x1, y2 - y1
- length = (dx**2 + dy**2) ** 0.5
- if length == 0:
- continue # Skip self-loops
- # Calculate perpendicular offset position
- mid_x = (x1 + x2) / 2 + (dy / length) * edge_label_offset
- mid_y = (y1 + y2) / 2 - (dx / length) * edge_label_offset
- ax.text(
- mid_x,
- mid_y,
- f"{weight:{edge_label_format}}",
- fontsize=edge_label_font_size,
- color=edge_label_color,
- ha="center",
- va="center",
- # bbox=dict(boxstyle="round", facecolor="white", alpha=0.8, edgecolor="none"),
- )
- # Edge weight colorbar
- if color_edges_by_weight and edge_colorbar and sm:
- cbar = plt.colorbar(sm, ax=ax, orientation="horizontal", shrink=0.3, aspect=40, pad=0.05)
- cbar.set_label("Edge Weight", fontsize=font_size)
- cbar.ax.tick_params(labelsize=font_size - 2)
- fig.suptitle(title, y=title_y, fontsize=font_size + 2, fontweight="bold", color="#333333")
- # Final layout adjustments
- plt.subplots_adjust(
- left=0.03, right=0.97, top=0.97, bottom=0.15 if (color_edges_by_weight and edge_colorbar) else 0.03
- )
- ax.set_axis_off()
- plt.margins(y=1) # Add 20% margin on all sides
- plt.savefig(f"/Users/kemalinecik/Downloads/fig2fig/{file_name}.pdf")
- # %%
- for trajectory_focus, result_df in result_df_main.groupby('trajectory'):
- if trajectory_focus == "Haem_ref":
- decompose_method = "2"
- input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{trajectory_focus}_litc_{decompose_method}.pkl")
- with open(input_trajectories_path_litc, 'rb') as _file:
- lit = pickle.load(_file)
- break
- df_edges = []
- for ind, row in result_df.iterrows():
- trj = lit.get_trajectory(row["booted_trajectory"], False)
- for from_cell, to_cell in trj.edges():
- row_copy = row.copy()
- row_copy["from_cell"] = from_cell
- row_copy["to_cell"] = to_cell
- df_edges.append(row_copy)
- df_edges = pd.DataFrame(df_edges)
- df_edges.reset_index(drop=True, inplace=True)
- # %%
- df_edges_pivot = df_edges.pivot(index=['path', 'booted_trajectory', 'trajectory', 'metric', 'from_cell', 'to_cell'], columns='representation', values='score')
- df_edges_pivot_rank = rank_metrics(df_edges_pivot, metric_direction_dict, metric_index_pos=3, method='min')
- df_edges_pivot_rank_mean = logistic_invert_rowwise(df_edges_pivot_rank).groupby(level=[4, 5]).mean()
- # %%
- from scipy.stats import mannwhitneyu, binomtest, wilcoxon
- def compare_scores_and_get_means(group):
- scpoli_scores = group['scpoli'].dropna()
- scvi_scores = group['scvi'].dropna()
- if len(scpoli_scores) < 2 or len(scvi_scores) < 2:
- raise ValueError
- # since your “scores” are actually ranks (ordinal) for the same boot replicate per edge, you have paired, ordinal data.
- # Perform the Mann-Whitney U test
- # stat, p_value = mannwhitneyu(scpoli_scores, scvi_scores, alternative='two-sided')
- # Wilcoxon signed-rank — more power if the distribution of paired differences is roughly symmetric; handle ties with zero_method="pratt"
- method = 'exact' if len(scpoli_scores) <= 25 else 'approx'
- p_value = wilcoxon(scpoli_scores, scvi_scores, zero_method='pratt', alternative='two-sided', method=method).pvalue
- # Statistically different: return means.
- return pd.Series({
- 'mean_scpoli': np.mean(scpoli_scores),
- 'mean_scvi': np.mean(scvi_scores),
- 'p_value': p_value
- })
- df_edges_pivot_rank_p = logistic_invert_rowwise(df_edges_pivot_rank).groupby(['from_cell', 'to_cell']).apply(compare_scores_and_get_means).reset_index()
- alpha = 0.05 / len(df_edges_pivot_rank_p)
- df_edges_pivot_rank_p['mean_scvi_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scvi'], np.nan)
- df_edges_pivot_rank_p['mean_scpoli_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scpoli'], np.nan)
- df_edges_pivot_rank_p
- # %%
- trj = InputTrajectories(ground_truth_trajectories).get_trajectory(trajectory_focus, False)
- representations = ["scvi", "scpoli"]
- for ind, row in df_edges_pivot_rank_p.iterrows():
- for repr in representations:
- trj[row["from_cell"]][row["to_cell"]][f"weight_{repr}"] = row[f"mean_{repr}_final"]
- # %%
- for i in ['weight_scpoli', 'weight_scvi']:
- plot_trajectory(
- trj,
- figsize=(4, 11),
- edge_weight_attribute=i,
- title="",
- edge_vmin=0.4,
- edge_vmax=0.7,
- color_edges_by_weight=True,
- node_size=150,
- node_color="lightgray",
- edge_color='black',
- edge_width=5,
- edge_labels=False,
- edge_label_font_size=5,
- edge_cmap='viridis',
- edge_colorbar=False,
- file_name=f"{trajectory_focus}_{i}"
- )
- # %%
- for trajectory_focus, result_df in result_df_main.groupby('trajectory'):
- if trajectory_focus == "germline":
- decompose_method = "2"
- input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{trajectory_focus}_litc_{decompose_method}.pkl")
- with open(input_trajectories_path_litc, 'rb') as _file:
- lit = pickle.load(_file)
- break
- df_edges = []
- for ind, row in result_df.iterrows():
- trj = lit.get_trajectory(row["booted_trajectory"], False)
- for from_cell, to_cell in trj.edges():
- row_copy = row.copy()
- row_copy["from_cell"] = from_cell
- row_copy["to_cell"] = to_cell
- df_edges.append(row_copy)
- df_edges = pd.DataFrame(df_edges)
- df_edges.reset_index(drop=True, inplace=True)
- df_edges_pivot = df_edges.pivot(index=['path', 'booted_trajectory', 'trajectory', 'metric', 'from_cell', 'to_cell'], columns='representation', values='score')
- df_edges_pivot_rank = rank_metrics(df_edges_pivot, metric_direction_dict, metric_index_pos=3, method='min')
- df_edges_pivot_rank_mean = logistic_invert_rowwise(df_edges_pivot_rank).groupby(level=[4, 5]).mean()
- df_edges_pivot_rank_p = logistic_invert_rowwise(df_edges_pivot_rank).groupby(['from_cell', 'to_cell']).apply(compare_scores_and_get_means).reset_index()
- alpha = 0.05 / len(df_edges_pivot_rank_p)
- df_edges_pivot_rank_p['mean_scvi_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scvi'], np.nan)
- df_edges_pivot_rank_p['mean_scpoli_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scpoli'], np.nan)
- trj = InputTrajectories(ground_truth_trajectories).get_trajectory(trajectory_focus, False)
- representations = ["scvi", "scpoli"]
- for ind, row in df_edges_pivot_rank_p.iterrows():
- for repr in representations:
- trj[row["from_cell"]][row["to_cell"]][f"weight_{repr}"] = row[f"mean_{repr}_final"]
- # %%
- for i in ['weight_scpoli', 'weight_scvi']:
- plot_trajectory(
- trj,
- figsize=(4, 0.75),
- edge_weight_attribute=i,
- title="",
- edge_vmin=0.4,
- edge_vmax=0.7,
- color_edges_by_weight=True,
- node_size=150,
- node_color="lightgray",
- edge_color='black',
- edge_width=5,
- edge_labels=False,
- edge_label_font_size=5,
- edge_cmap='viridis',
- edge_colorbar=False,
- file_name=f"{trajectory_focus}_{i}"
- )
- # %%
- def compare_scores_and_get_means(group):
- tardis_scores = group['tardis'].dropna()
- scvi_scores = group['scvi'].dropna()
- if len(tardis_scores) < 2 or len(scvi_scores) < 2:
- raise ValueError
- # since your “scores” are actually ranks (ordinal) for the same boot replicate per edge, you have paired, ordinal data.
- # Perform the Mann-Whitney U test
- # stat, p_value = mannwhitneyu(tardis_scores, scvi_scores, alternative='two-sided')
- # Wilcoxon signed-rank — more power if the distribution of paired differences is roughly symmetric; handle ties with zero_method="pratt"
- method = 'exact' if len(tardis_scores) <= 25 else 'approx'
- p_value = wilcoxon(tardis_scores, scvi_scores, zero_method='pratt', alternative='two-sided', method=method).pvalue
- # Statistically different: return means.
- return pd.Series({
- 'mean_tardis': np.mean(tardis_scores),
- 'mean_scvi': np.mean(scvi_scores),
- 'p_value': p_value
- })
- # %%
- for trajectory_focus, result_df in result_df_main.groupby('trajectory'):
- if trajectory_focus == "pns_neuro":
- decompose_method = "2"
- input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{trajectory_focus}_litc_{decompose_method}.pkl")
- with open(input_trajectories_path_litc, 'rb') as _file:
- lit = pickle.load(_file)
- break
- df_edges = []
- for ind, row in result_df.iterrows():
- trj = lit.get_trajectory(row["booted_trajectory"], False)
- for from_cell, to_cell in trj.edges():
- row_copy = row.copy()
- row_copy["from_cell"] = from_cell
- row_copy["to_cell"] = to_cell
- df_edges.append(row_copy)
- df_edges = pd.DataFrame(df_edges)
- df_edges.reset_index(drop=True, inplace=True)
- df_edges_pivot = df_edges.pivot(index=['path', 'booted_trajectory', 'trajectory', 'metric', 'from_cell', 'to_cell'], columns='representation', values='score')
- df_edges_pivot_rank = rank_metrics(df_edges_pivot, metric_direction_dict, metric_index_pos=3, method='min')
- df_edges_pivot_rank_mean = logistic_invert_rowwise(df_edges_pivot_rank).groupby(level=[4, 5]).mean()
- df_edges_pivot_rank_p = logistic_invert_rowwise(df_edges_pivot_rank).groupby(['from_cell', 'to_cell']).apply(compare_scores_and_get_means).reset_index()
- alpha = 0.05 / len(df_edges_pivot_rank_p)
- df_edges_pivot_rank_p['mean_scvi_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_scvi'], np.nan)
- df_edges_pivot_rank_p['mean_tardis_final'] = np.where(df_edges_pivot_rank_p['p_value'] < alpha, df_edges_pivot_rank_p['mean_tardis'], np.nan)
- trj = InputTrajectories(ground_truth_trajectories).get_trajectory(trajectory_focus, False)
- representations = ["scvi", "tardis"]
- for ind, row in df_edges_pivot_rank_p.iterrows():
- for repr in representations:
- trj[row["from_cell"]][row["to_cell"]][f"weight_{repr}"] = row[f"mean_{repr}_final"]
- # %%
- for i in ['weight_tardis', 'weight_scvi']:
- plot_trajectory(
- trj,
- figsize=(4, 3.25),
- edge_weight_attribute=i,
- title="",
- edge_vmin=0.4,
- edge_vmax=0.7,
- color_edges_by_weight=True,
- node_size=150,
- node_color="lightgray",
- edge_color='black',
- edge_width=5,
- edge_labels=False,
- edge_label_font_size=5,
- edge_cmap='viridis',
- edge_colorbar=False,
- file_name=f"{trajectory_focus}_{i}"
- )
- # %%
- from matplotlib import ticker
- def create_vertical_colorbar(
- label,
- filename,
- vmin=2.5,
- vmax=3.5,
- cmap='plasma',
- fontsize=10,
- n_ticks=5,
- fig_height=4,
- fig_width=1,
- scale=1.0
- ):
- """
- Creates and displays a standalone vertical colorbar.
- Parameters:
- - label: str, label for the colorbar
- - vmin, vmax: float, data range
- - cmap: str, colormap name
- - fontsize: int, for label & tick labels
- - n_ticks: int, approx number of ticks to show
- - fig_height: float, base height in inches
- - fig_width: float, base width in inches
- - scale: float, scale factor for overall size
- """
- # 1) Make the ScalarMappable
- norm = mpl.colors.Normalize(vmin=vmin, vmax=vmax)
- sm = mpl.cm.ScalarMappable(norm=norm, cmap=cmap)
- sm.set_array([])
- # 2) New figure, scaled
- fig = plt.figure(figsize=(fig_width * scale, fig_height * scale))
- # you can tweak these numbers if you want more/less padding
- cax = fig.add_axes([0.2, 0.05, 0.1, 0.9])
- # 3) Draw the colorbar
- cbar = fig.colorbar(
- sm,
- cax=cax,
- orientation='vertical'
- )
- # 4) Automatic “nice” locators + re‐draw
- cbar.locator = ticker.MaxNLocator(nbins=n_ticks)
- cbar.update_ticks()
- # 5) Force one‐decimal formatting
- cbar.ax.yaxis.set_major_formatter(ticker.FormatStrFormatter('%.1f'))
- # 6) Label & style
- cbar.set_label(label, fontsize=fontsize)
- cbar.ax.tick_params(labelsize=fontsize)
- plt.subplots_adjust(left=0.03, right=0.97, top=0.97, bottom=0.15)
- plt.savefig(f"/Users/kemalinecik/Downloads/fig2fig/{filename}.pdf")
- create_vertical_colorbar("Performance", filename = "colorbar_performance_rank",
- cmap="viridis", n_ticks=3, fig_height=2, fig_width=2.4, scale=0.8, fontsize=10, vmin=0.4, vmax=0.7)
- # %% [markdown]
- # # Figure 3
- # %%
- result_df_boot = pd.read_pickle(os.path.join(temp_move, f"df_boots_fig2_calculations.pkl"))
- result_df_boot = result_df_boot[result_df_boot["representation"]=="scvi"]
- result_df_boot.reset_index(inplace=True, drop=True)
- result_df_main = pd.read_pickle(os.path.join(temp_move, f"df_boots_fig2_calculations_for_individuals.pkl"))
- # %%
- focus_trajectory = "germline"
- df_compo = result_df_main[result_df_main["trajectory"] == focus_trajectory]
- df_atlas = result_df_boot[result_df_boot["trajectory"] == focus_trajectory]
- df_atlas.reset_index(inplace=True, drop=True)
- df_compo.reset_index(inplace=True, drop=True)
- decompose_method = "2"
- input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{focus_trajectory}_litc_{decompose_method}.pkl")
- with open(input_trajectories_path_litc, 'rb') as _file:
- lit = pickle.load(_file)
- trj_size_dict = dict()
- for trj_name in lit.graph['trajectories']:
- trj_size_dict[trj_name] = len(lit.get_trajectory(trj_name, False).nodes()), len(lit.get_trajectory(trj_name, False).edges())
- assert df_compo.shape == df_atlas.shape
- for i in range(len(df_compo)):
- i1 = list(df_atlas.iloc[i][["booted_trajectory", "metric"]])
- i2 = list(df_compo.iloc[i][["booted_trajectory", "metric"]])
- assert i1==i2
- # df_both
- df_both = df_compo.copy()
- del df_both["score"]
- df_both["score_Component"] = list(df_compo["score"])
- df_both["score_Atlas"] = list(df_atlas["score"])
- df_both.reset_index(inplace=True, drop=True)
- del df_both["anndata_subset"]
- def get_winners(row):
- metric = row['metric']
- direction = metric_direction_dict[metric]
- c = row['score_Component']
- a = row['score_Atlas']
- if direction == 'decreasing':
- if c > a:
- return pd.Series([0, 1])
- elif a > c:
- return pd.Series([1, 0])
- else:
- return pd.Series([0.5, 0.5])
- elif direction == 'increasing':
- if c < a:
- return pd.Series([0, 1])
- elif a < c:
- return pd.Series([1, 0])
- else:
- return pd.Series([0.5, 0.5])
- else:
- return pd.Series([None, None]) # In case metric not found
- df_both[['winner_Component', 'winner_Atlas']] = df_both.apply(get_winners, axis=1)
- def get_nodes_edges(row):
- trj_name = row['booted_trajectory']
- len_nodes, len_edges = trj_size_dict[trj_name]
- return pd.Series([len_nodes, len_edges])
- df_both[['len_nodes', 'len_edges']] = df_both.apply(get_nodes_edges, axis=1)
- label_mapping_o = {
- "adjacency": "Adjacency",
- "embedding": "Embedding",
- "pseudotime": "Pseudotime",
- }
- df_both["path"] = [label_mapping_o[i] for i in df_both["path"]]
- # long_both
- id_cols = [
- "decompose_method", "trajectory_class", "trajectory",
- "booted_trajectory", 'len_nodes', 'len_edges', "path", "metric",
- "trajectory_annotation_column"
- ]
- _df = df_both.reset_index()
- id_cols = ["index"] + id_cols
- long_both = (
- pd.wide_to_long(
- _df, # the original wide dataframe
- stubnames=["score", "winner"],
- i=id_cols, # columns that make a unique row
- j="focus", # new column: 'component' or 'atlas'
- sep="_", # underscore separates stub & suffix
- suffix="(Component|Atlas)" # the two suffix values to capture
- )
- .reset_index()
- .sort_values(id_cols + ["focus"]) # optional tidy ordering
- )
- del long_both["index"]
- replacements = {
- 'Pseudotime': 'Cell Order\ncontinuity',
- 'Embedding': 'Trajectory\ndominance',
- 'Adjacency': 'Topology\nconservation'
- }
- long_both_p = long_both.pivot(index=["booted_trajectory", "path", "metric"], columns='focus', values='winner')
- long_both_p = long_both_p.groupby(level=[0, 1]).mean()
- long_both_p.rename(index=replacements, level="path", inplace=True)
- long_both_back = (
- long_both_p
- .stack(dropna=False) # <- dropna=True by default
- .rename('winner')
- .reset_index() # gives: booted_trajectory, path, metric, focus, winner
- )
- long_both_back.head()
- # %%
- label_mapping = {
- "Atlas": "IA",
- "Component": "Garcia",
- }
- # match boxplot’s height & per-column width
- fig_height = 3
- n_cols = long_both_back["focus"].nunique()
- col_width = 2.4
- fig_width = n_cols * col_width
- fig, ax = plt.subplots(figsize=(fig_width, fig_height))
- # draw bars
- bp = sns.barplot(
- data=long_both_back,
- hue="focus",
- y="winner",
- x="path",
- palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
- width=0.8,
- saturation=1,
- errorbar="sd",
- capsize=0.05,
- err_kws={'color': 'black', 'linewidth': 0.81},
- edgecolor="black",
- linewidth=0.81,
- zorder=1,
- ax=ax
- )
- # enforce edge styling
- for patch in bp.patches:
- patch.set_edgecolor("black")
- patch.set_linewidth(0.81)
- # axes
- ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
- ax.tick_params(axis="y", labelsize=10)
- ax.set_xlabel('Germline', fontsize=10, labelpad=10)
- plt.ylabel("Performance", fontsize=10, labelpad=10)
- ax.set_ylim(0, None)
- ax.yaxis.set_major_formatter(FormatStrFormatter('%.1f'))
- # grid + despine
- ax.grid(axis='y', linestyle='--', alpha=0.4)
- sns.despine(left=True, trim=False)
- # remove the old legend
- ax.legend_.remove()
- # build new, professional legend
- handles, labels = ax.get_legend_handles_labels()
- new_labels = [label_mapping.get(l, l) for l in labels]
- slim_handles = [
- mpatches.Patch(
- facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
- edgecolor="black",
- linewidth=0.5 # thinner border
- )
- for h in handles
- ]
- leg = ax.legend(
- slim_handles,
- new_labels,
- title="Representation",
- loc="center left",
- bbox_to_anchor=(1.02, 0.5),
- alignment="left",
- frameon=False,
- title_fontsize=8,
- fontsize=8,
- borderaxespad=0.5,
- handletextpad=0.4,
- labelspacing=0.5,
- )
- # tighten layout (accounts for legend)
- plt.tight_layout(rect=[0, 0, 0.85, 1])
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/comparison_germline_component.pdf")
- # %%
- long_both_germ = long_both_back.copy()
- # %%
- focus_trajectory = "Haem_ref"
- df_compo = result_df_main[result_df_main["trajectory"] == focus_trajectory]
- df_atlas = result_df_boot[result_df_boot["trajectory"] == focus_trajectory]
- df_atlas.reset_index(inplace=True, drop=True)
- df_compo.reset_index(inplace=True, drop=True)
- decompose_method = "2"
- input_trajectories_path_litc = os.path.join(temp_move, f"adata_hdca_input_{focus_trajectory}_litc_{decompose_method}.pkl")
- with open(input_trajectories_path_litc, 'rb') as _file:
- lit = pickle.load(_file)
- trj_size_dict = dict()
- for trj_name in lit.graph['trajectories']:
- trj_size_dict[trj_name] = len(lit.get_trajectory(trj_name, False).nodes()), len(lit.get_trajectory(trj_name, False).edges())
- assert df_compo.shape == df_atlas.shape
- for i in range(len(df_compo)):
- i1 = list(df_atlas.iloc[i][["booted_trajectory", "metric"]])
- i2 = list(df_compo.iloc[i][["booted_trajectory", "metric"]])
- assert i1==i2
- # df_both
- df_both = df_compo.copy()
- del df_both["score"]
- df_both["score_Component"] = list(df_compo["score"])
- df_both["score_Atlas"] = list(df_atlas["score"])
- df_both.reset_index(inplace=True, drop=True)
- del df_both["anndata_subset"]
- def get_winners(row):
- metric = row['metric']
- direction = metric_direction_dict[metric]
- c = row['score_Component']
- a = row['score_Atlas']
- if direction == 'decreasing':
- if c > a:
- return pd.Series([0, 1])
- elif a > c:
- return pd.Series([1, 0])
- else:
- return pd.Series([0.5, 0.5])
- elif direction == 'increasing':
- if c < a:
- return pd.Series([0, 1])
- elif a < c:
- return pd.Series([1, 0])
- else:
- return pd.Series([0.5, 0.5])
- else:
- return pd.Series([None, None]) # In case metric not found
- df_both[['winner_Component', 'winner_Atlas']] = df_both.apply(get_winners, axis=1)
- def get_nodes_edges(row):
- trj_name = row['booted_trajectory']
- len_nodes, len_edges = trj_size_dict[trj_name]
- return pd.Series([len_nodes, len_edges])
- df_both[['len_nodes', 'len_edges']] = df_both.apply(get_nodes_edges, axis=1)
- label_mapping_o = {
- "adjacency": "Adjacency",
- "embedding": "Embedding",
- "pseudotime": "Pseudotime",
- }
- df_both["path"] = [label_mapping_o[i] for i in df_both["path"]]
- # long_both
- id_cols = [
- "decompose_method", "trajectory_class", "trajectory",
- "booted_trajectory", 'len_nodes', 'len_edges', "path", "metric",
- "trajectory_annotation_column"
- ]
- _df = df_both.reset_index()
- id_cols = ["index"] + id_cols
- long_both = (
- pd.wide_to_long(
- _df, # the original wide dataframe
- stubnames=["score", "winner"],
- i=id_cols, # columns that make a unique row
- j="focus", # new column: 'component' or 'atlas'
- sep="_", # underscore separates stub & suffix
- suffix="(Component|Atlas)" # the two suffix values to capture
- )
- .reset_index()
- .sort_values(id_cols + ["focus"]) # optional tidy ordering
- )
- del long_both["index"]
- replacements = {
- 'Pseudotime': 'Cell Order\ncontinuity',
- 'Embedding': 'Trajectory\ndominance',
- 'Adjacency': 'Topology\nconservation'
- }
- long_both_p = long_both.pivot(index=["booted_trajectory", "path", "metric"], columns='focus', values='winner')
- long_both_p = long_both_p.groupby(level=[0, 1]).mean()
- long_both_p.rename(index=replacements, level="path", inplace=True)
- long_both_back = (
- long_both_p
- .stack(dropna=False) # <- dropna=True by default
- .rename('winner')
- .reset_index() # gives: booted_trajectory, path, metric, focus, winner
- )
- long_both_back.head()
- # %%
- label_mapping = {
- "Atlas": "IA",
- "Component": "Suo",
- }
- # match boxplot’s height & per-column width
- fig_height = 3
- n_cols = long_both_back["focus"].nunique()
- col_width = 2.4
- fig_width = n_cols * col_width
- fig, ax = plt.subplots(figsize=(fig_width, fig_height))
- # draw bars
- bp = sns.barplot(
- data=long_both_back,
- hue="focus",
- y="winner",
- x="path",
- palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
- width=0.8,
- saturation=1,
- errorbar="sd",
- capsize=0.05,
- err_kws={'color': 'black', 'linewidth': 0.81},
- edgecolor="black",
- linewidth=0.81,
- zorder=1,
- ax=ax
- )
- # enforce edge styling
- for patch in bp.patches:
- patch.set_edgecolor("black")
- patch.set_linewidth(0.81)
- # axes
- ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
- ax.tick_params(axis="y", labelsize=10)
- ax.set_xlabel('Hematopoietic Lineage', fontsize=10, labelpad=10)
- plt.ylabel("Performance", fontsize=10, labelpad=10)
- ax.set_ylim(0, None)
- ax.yaxis.set_major_formatter(FormatStrFormatter('%.1f'))
- # grid + despine
- ax.grid(axis='y', linestyle='--', alpha=0.4)
- sns.despine(left=True, trim=False)
- # remove the old legend
- ax.legend_.remove()
- # build new, professional legend
- handles, labels = ax.get_legend_handles_labels()
- new_labels = [label_mapping.get(l, l) for l in labels]
- slim_handles = [
- mpatches.Patch(
- facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
- edgecolor="black",
- linewidth=0.5 # thinner border
- )
- for h in handles
- ]
- leg = ax.legend(
- slim_handles,
- new_labels,
- title="Representation",
- loc="center left",
- bbox_to_anchor=(1.02, 0.5),
- alignment="left",
- frameon=False,
- title_fontsize=8,
- fontsize=8,
- borderaxespad=0.5,
- handletextpad=0.4,
- labelspacing=0.5,
- )
- # tighten layout (accounts for legend)
- plt.tight_layout(rect=[0, 0, 0.85, 1])
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/comparison_hemato_component.pdf")
- # %%
- long_both_haem = long_both_back.copy()
- # %%
- long_both_haem["trajectory_class"] = "Hematopoietic\nLineage"
- long_both_germ["trajectory_class"] = "Germline"
- long_both2 = pd.concat([long_both_haem, long_both_germ])
- long_both2.reset_index(drop=True, inplace=True)
- # %%
- label_mapping = {
- "Atlas": "Atlas",
- "Component": "Component",
- }
- # match boxplot’s height & per-column width
- fig_height = 3
- n_cols = long_both2["focus"].nunique()
- col_width = 2
- fig_width = n_cols * col_width
- fig, ax = plt.subplots(figsize=(fig_width, fig_height))
- # draw bars
- bp = sns.barplot(
- data=long_both2,
- hue="focus",
- y="winner",
- x="trajectory_class",
- palette=sns.color_palette("Greys", n_colors=len(label_mapping)+2),
- width=0.8,
- saturation=1,
- errorbar="sd",
- capsize=0.05,
- err_kws={'color': 'black', 'linewidth': 0.81},
- edgecolor="black",
- linewidth=0.81,
- zorder=1,
- ax=ax
- )
- # enforce edge styling
- for patch in bp.patches:
- patch.set_edgecolor("black")
- patch.set_linewidth(0.81)
- # axes
- ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center', fontsize=10)
- ax.tick_params(axis="y", labelsize=10)
- ax.set_xlabel('Lineage', fontsize=10, labelpad=10)
- plt.ylabel("Performance", fontsize=10, labelpad=10)
- ax.set_ylim(0, None)
- ax.yaxis.set_major_formatter(FormatStrFormatter('%.1f'))
- # grid + despine
- ax.grid(axis='y', linestyle='--', alpha=0.4)
- sns.despine(left=True, trim=False)
- # remove the old legend
- ax.legend_.remove()
- # build new, professional legend
- handles, labels = ax.get_legend_handles_labels()
- new_labels = [label_mapping.get(l, l) for l in labels]
- slim_handles = [
- mpatches.Patch(
- facecolor=h.get_facecolor(), # <-- keep the whole RGBA tuple
- edgecolor="black",
- linewidth=0.5 # thinner border
- )
- for h in handles
- ]
- leg = ax.legend(
- slim_handles,
- new_labels,
- title="Representation",
- loc="center left",
- bbox_to_anchor=(1.02, 0.5),
- alignment="left",
- frameon=False,
- title_fontsize=8,
- fontsize=8,
- borderaxespad=0.5,
- handletextpad=0.4,
- labelspacing=0.5,
- )
- # tighten layout (accounts for legend)
- plt.tight_layout(rect=[0, 0, 0.85, 1])
- plt.savefig("/Users/kemalinecik/Downloads/fig2fig/comparison_germline_hemato_both.pdf")
- # %%
- 1
- # %% [markdown]
- # # Figure UMAPs
- # %%
- import scvelo as scv
- import cellrank as cr
- import scanpy.external as sce
- scv.settings.verbosity = 3
- cr.settings.verbosity = 2
- sc.settings.verbose = 3
- # %%
- adata_comp = ad.read_h5ad(os.path.join(temp_move, f"anndata_hdca_input_germline_litc_comparison_comp.h5ad"))
- adata_atlas = ad.read_h5ad(os.path.join(temp_move, f"anndata_hdca_input_germline_litc_comparison_atlas.h5ad"))
- # %%
- display((adata_atlas.shape, adata_comp.shape))
- sc.pp.neighbors(adata_atlas, n_neighbors=15)
- sc.tl.umap(adata_atlas)
- sc.pp.neighbors(adata_comp, n_neighbors=15)
- sc.tl.umap(adata_comp)
- # %%
- (
- df_hdca_component_adata_paths,
- df_hdca_lvl3_adata_paths,
- out_df,
- ground_truth_trajectories,
- ground_truth_trajectories_annotation_level,
- trajectory_subset_dict
- ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
- # %%
- # Mapping from original to formal names
- celltype_mapping = {
- "GC": "Granulosa Cell",
- "PGC": "Primordial Germ Cell",
- "oocyte": "Oocyte",
- "oogonia": "Oogonia",
- "oogonia_meiotic_1": "Oogonia\n(Meiotic I)",
- "oogonia_meiotic_2": "Oogonia\n(Meiotic II)",
- "pre_oocyte": "Pre-Oocyte",
- "pre_spermatogonia": "Pre-Spermatogonia"
- }
- # Apply mapping to the annotations
- adata_atlas.obs[ground_truth_trajectories_annotation_level[focus_trajectory]] = \
- adata_atlas.obs[ground_truth_trajectories_annotation_level[focus_trajectory]].map(celltype_mapping)
- # Apply mapping to the annotations
- adata_comp.obs[ground_truth_trajectories_annotation_level[focus_trajectory]] = \
- adata_comp.obs[ground_truth_trajectories_annotation_level[focus_trajectory]].map(celltype_mapping)
- # %%
- def cr_calc(adata_cr, root_cell_type, i, ann_levl, knn=15):
- sc.tl.pca(adata_cr, n_comps=11, svd_solver='arpack')
- sce.tl.palantir(adata_cr, n_components=10, knn=knn)
- root_cells = adata_cr.obs_names[adata_cr.obs[ann_levl] == root_cell_type]
- if len(root_cells) == 0:
- raise ValueError(f"No cells found with {ann_levl!r} == {root_cell_type!r}")
- start_cell = root_cells[i]
- pr_res = sce.tl.palantir_results(
- adata_cr,
- early_cell=start_cell,
- ms_data="X_palantir_multiscale",
- knn=knn,
- num_waypoints=1200,
- n_jobs=-1
- )
- adata_cr.obs["palantir_pseudotime"] = pr_res.pseudotime
- pk = cr.kernels.PseudotimeKernel(adata_cr, time_key='palantir_pseudotime')
- pk.compute_transition_matrix()
- return pk
- # %%
- focus_trajectory = "germline"
- ann_lvl = ground_truth_trajectories_annotation_level[focus_trajectory]
- pk_adata_atlas = cr_calc(adata_atlas.copy(), root_cell_type = "Primordial Germ Cell", i=10, ann_levl=ann_lvl)
- # %%
- plt.figure(figsize=(4,3))
- pk_adata_atlas.plot_projection(basis='umap', recompute=False, color=[ann_lvl], title="", ax=plt.gca(), legend_fontsize="medium", legend_fontweight="bold", legend_fontoutline=2,
- save="/Users/kemalinecik/Downloads/fig2fig/comparison_germline_atlas_umap.pdf")
- # %%
- pk_adata_comp = cr_calc(adata_comp.copy(), root_cell_type = "Primordial Germ Cell", i=10, ann_levl=ann_lvl)
- # %%
- plt.figure(figsize=(4,3))
- pk_adata_comp.plot_projection(basis='umap', recompute=False, color=[ann_lvl], title="", ax=plt.gca(), legend_fontsize="medium", legend_fontweight="bold", legend_fontoutline=2,
- save="/Users/kemalinecik/Downloads/fig2fig/comparison_germline_component_umap.pdf")
- # %% [markdown]
- # # Figure UMAPs
- # %%
- import scvelo as scv
- import cellrank as cr
- import scanpy.external as sce
- scv.settings.verbosity = 3
- cr.settings.verbosity = 2
- sc.settings.verbose = 3
- (
- df_hdca_component_adata_paths,
- df_hdca_lvl3_adata_paths,
- out_df,
- ground_truth_trajectories,
- ground_truth_trajectories_annotation_level,
- trajectory_subset_dict
- ) = load_pickle_bundle(os.path.join(temp_move, "_experiment_development_atlas_helper_pickle_bundle.pkl"))
- # %%
- def cr_calc(adata_cr, root_cell_type, i, ann_levl, knn=15):
- sc.tl.pca(adata_cr, n_comps=11, svd_solver='arpack')
- sce.tl.palantir(adata_cr, n_components=10, knn=knn)
- root_cells = adata_cr.obs_names[adata_cr.obs[ann_levl] == root_cell_type]
- if len(root_cells) == 0:
- raise ValueError(f"No cells found with {ann_levl!r} == {root_cell_type!r}")
- start_cell = root_cells[i]
- pr_res = sce.tl.palantir_results(
- adata_cr,
- early_cell=start_cell,
- ms_data="X_palantir_multiscale",
- knn=knn,
- num_waypoints=1200,
- n_jobs=-1
- )
- adata_cr.obs["palantir_pseudotime"] = pr_res.pseudotime
- pk = cr.kernels.PseudotimeKernel(adata_cr, time_key='palantir_pseudotime')
- pk.compute_transition_matrix()
- return pk
- # %%
- # Mapping for B-cell and progenitor annotations
- bcell_mapping = {
- "Mature_B": "Mature B",
- "Large_pre_B": "Large Pre B",
- "Small_pre_B": "Small Pre B",
- "B1": "B1",
- "Pro_B": "Pro B",
- "Late_Pro_B": "Late Pro B",
- "HSC_MPP": "HSC MPP",
- "Immature_B": "Immature B",
- "Pre_pro_B": "Pre-pro B",
- "LMPP_MLP": "LMPP MLP",
- "CYCLING_B": "Cycling B",
- "Plasma_B": "Plasma B"
- }
- # %%
- for use_rep in ['scpoli', 'scvi']:
- print("\n\n\n\%%%%%%%%%%%%%%\n\n\n", use_rep)
- adata_atlas = ad.read_h5ad(os.path.join(temp_move, f"anndata_hdca_input_Haem_ref_litc_2_{use_rep}_nodes_b.h5ad"))
- print(adata_atlas.shape)
- sc.pp.subsample(adata_atlas, fraction=0.2)
- print(adata_atlas.shape)
- sc.pp.neighbors(adata_atlas, n_neighbors=15)
- sc.tl.umap(adata_atlas)
- ann_lvl = ground_truth_trajectories_annotation_level["Haem_ref"]
- adata_atlas.obs[ann_lvl] = adata_atlas.obs[ann_lvl].map(bcell_mapping)
- pk_adata_atlas = cr_calc(adata_atlas.copy(), root_cell_type = "HSC MPP", i=10, ann_levl=ann_lvl, knn=30)
- plt.figure(figsize=(5,4))
- pk_adata_atlas.plot_projection(basis='umap', recompute=False, color=[ann_lvl], title="", ax=plt.gca(), legend_fontsize="medium", legend_fontweight="bold", legend_fontoutline=2,
- save=f"/Users/kemalinecik/Downloads/fig2fig/comparison_models_b_umap_{use_rep}.svg")
- # break
- del adata_atlas, pk_adata_atlas
- gc.collect()
- # %%
sctram_figures_only.ipynb at commit e570de5, under BSD-3-Clause · at the source
Overview
and 74 other authors
Julia Karjalainen9, Anna Laddach10, Shaista Madad2, John EG Lawrence2,8, Vitalii Kleshchevnikov2, Steven Lisgo1, Jimmy Tsz Hang Lee2, Joséphine Blevinal11, Ahlam Alqahtani1, Stanislaw Makarchuk2, Jake Jackson1, Ekin Ucuncu5, Teresa P Silva5, Valentina Lorenzi2,6, Fereshteh Torabi2, Rachel A Botting1, Kenny Roberts2, Bayanne Olabi1, Keerthi P Chakala2, Leander Dony4,12, Gabriele Dall’Aglio4, Ana-Maria Cujba3, Holly J Whitfield2, Christine Hale2, Justin Engelbert1, Tong Li2, Elena Prigmore2, Nusayhah Hudaa Gopee1,2, Michael W Mather1, Ben Talks1, Lijiang Fei3, Diana Pereira13, Corina Moldovan14, Rasa Elmentaite3, Krzysztof Polanski3, Alexander V Predeus2, Pasha Mazin2, Vicky Rowe2, Semih Bayraktar3, Vincent R Knight-Schrijver3, Menna Clatworthy15, Kazumasa Kanemaru3, James Cranley3,16, Andrew Guo15, Sanjay Sinha3, David McDonald1, Andrew Filby1, Janet Kerwin1, Peng He17,18, Inês Sequeira13, Roser Vento-Tormo2, Malte D. Luecken4,19, Alain Chédotal11,20,21, Deanne Taylor22,23, Silvia Martin-Almedina24, Pia Ostergaard24, Sam Behjati2,25, Elizabeth Robertson26, Nicola K Wilson3,7, John Marioni5,27, Helen V Firth2,28, Caroline F Wright29, Olivier Pourquie30, Laure Gambardella2, Paula Alexandre5, Berthold Gottgens3,7, Kerstin B Meyer2, MS Vijayabaskar2, Mo Lotfollahi2, Sarah A Teichmann3,16,31, Fabian J Theis4, Fränze Progatzky9, Omer Ali Bayraktar2, Muzlifah Haniffa2,28,3232 affiliations
- Biosciences Institute, Newcastle University, Newcastle upon Tyne, UK
- Wellcome Sanger Institute, Wellcome Genome Campus, Hinxton, UK
- Cambridge Stem Cell Institute, Jeffrey Cheah Biomedical Centre, Cambridge Biomedical Campus, University of Cambridge, Cambridge, UK
- Institute of Computational Biology, Computational Health Center, Helmholtz Munich, Neuherberg, Germany
- Developmental Biology & Cancer Department, University College London (UCL) Great Ormond Street Institute of Child Health, UCL, London, UK
- European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Cambridge, UK
- Department of Haematology, University of Cambridge, Cambridge, UK
- Department of Surgery, University of Cambridge, Cambridge, UK
- Kennedy Institute of Rheumatology, University of Oxford, Oxford, UK
- Nervous System Development and Homeostasis Laboratory, the Francis Crick Institute, London, UK
- Sorbonne Université, INSERM, CNRS, Institut de la Vision, Paris, France
- Department of Biosystems Science and Engineering, ETH Zürich, Basel, Switzerland
- Center for Oral Immunobiology and Regenerative Medicine, Barts Centre for Squamous Cancer, Institute of Dentistry, Queen Mary University of London, London, UK
- Department of Pathology, Newcastle Hospitals NHS Foundation Trust, Newcastle upon Tyne, UK
- Molecular Immunity Unit, Department of Medicine, University of Cambridge, Cambridge, UK
- Department of Medicine, University of Cambridge, Cambridge, UK
- Department of Pathology, University of California, San Francisco, CA, USA
- Eli and Edythe Broad Center of Regeneration Medicine and Stem Cell Research, University of California, San Francisco, San Francisco, CA, USA
- Institute of Lung Health and Immunity (LHI), Helmholtz Munich, Comprehensive Pneumology Center (CPC-M), Germany; Member of the German Center for Lung Research (DZL)
- Institut de Pathologie, Groupe Hospitalier Est, Hospices Civils de Lyon, Lyon, France
- University Claude Bernard Lyon 1, MeLiS, CNRS UMR5284, INSERM U1314, 69008, Lyon, France
- Department of Biomedical and Health Informatics (DBHi), The Children’s Hospital of Philadelphia, Philadelphia, PA, USA
- Department of Pediatrics, University of Pennsylvania Perelman School of Medicine, Philadelphia, PA, USA
- School of Health and Medical Sciences, City St George’s University of London, London, UK
- Department of Paediatrics, University of Cambridge, Cambridge, UK
- Sir William Dunn School of Pathology, University of Oxford, Oxford, UK
- Cancer Research UK, Cambridge Institute, University of Cambridge, Cambridge, UK
- Cambridge University Hospitals Foundation Trust, Addenbrooke’s Hospital, Cambridge, UK
- Department of Clinical and Biomedical Sciences, St Luke’s Campus, University of Exeter, Exeter, UK
- Department of Genetics, Harvard Medical School and Department of Pathology, Brigham and Women’s Hospital, Boston, MA, USA
- CIFAR Macmillan Multi-scale Human Programme, CIFAR, Toronto, Canada
- Cambridge Institute for Therapeutic Immunology and Infectious Diseases (CIITID), University of Cambridge, UK
Abstract
Single cell genomics has enabled analysis of human prenatal development at unprecedented resolution. However, most studies have relied on dissociated tissues during restricted windows of development, limiting insights into how spatially distributed networks of cells, and multicellular niches emerge and adapt to distinct organ microenvironments in situ. Moreover, existing human developmental atlases have not yet been harmonised, and we thus lack a comprehensive catalogue of known cell types in the developing human body.
Here, we introduce the Human Developmental Cell Atlas (HDCA), a unified structural, cellular and molecular resource for prenatal human development. The HDCA integrates published and unpublished single cell/
For a global overview of the human embryo’s multicellular communities, we applied unsupervised deep learning to our intact human embryo spatial data, charting 114 tissue niches that are structural and signalling hubs for the cellular interactions of the embryo. Guided by these niches, we profiled cellular networks over space and time, not examinable using single-organ atlases. In so doing, we revealed tissue-specific fibroblast patterning from previously undescribed mesenchyme progenitors, early diversification of organ-specific blood capillaries and lymphatic vasculature, emergence of neural crest cell fates, the formation of placode- and neural crest-derived peripheral sensory neurons, and how tissue niches guide peripheral neuron maturation and axonal migration. The HDCA thus serves as a comprehensive step towards a comprehensive understanding of human prenatal development, and a template towards unravelling the biology of congenital disorders.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 19 matches between paragraphs and lines of code.
haniffalab/hdca-paper
Availability: 1 check, the latest on 28 September 2026: the link is dead
- 28 September 2026: the link is dead
theislab/idtrack
d8514ab51a494842f4278e769c9538f103592c7a, 20 January 2026Availability: 1 check, the latest on 28 September 2026: the link answers
- 28 September 2026: the link answers
63 files
- docs/
_notebooks/ , Jupyter, 264 lines00_idtrack_overview.ipyn b - docs/
_notebooks/ , Jupyter, 187 lines01_installation_guide.ip ynb - docs/
_notebooks/ , Jupyter, 627 lines02_prepare_new_external_ yaml.ipynb - docs/
_notebooks/ , Jupyter, 289 lines03_initialization_graph. ipynb - docs/
_notebooks/ , Jupyter, 307 lines04_api_deep_dive_human.i pynb - docs/
_notebooks/ , Jupyter, 575 lines05_tutorial_harmonizatio n.ipynb - docs/
_notebooks/ , Jupyter, 240 lines06_tutorial_humanization _mouse_pig_to_human.ipyn b - docs/
_notebooks/ , Jupyter, 431 lines07_advanced_topics.ipynb - docs/
_notebooks/ , Jupyter, 185 lines99_initialization_test.i pynb - docs/
_notebooks/ , Python, 381 lines_notebook_utils.py - docs/
conf.py , Python, 215 lines - example_manual_running copy.ipynb, Jupyter, 194 lines
- idtrack/
__init__.py , Python, 51 lines - idtrack/
__main__.py , Python, 125 lines - idtrack/
_api.py , Python, 697 lines - idtrack/
_connection_bridge.py , Python, 431 lines - idtrack/
_database_manager.py , Python, 3,045 lines - idtrack/
_db.py , Python, 420 lines - idtrack/
_external_databases.py , Python, 385 lines - idtrack/
_external_mappers/ , Python, 64 lines__init__.py - idtrack/
_external_mappers/ , Python, 270 lines_backend_gget.py - idtrack/
_external_mappers/ , Python, 398 lines_backend_gprofiler.py - idtrack/
_external_mappers/ , Python, 277 lines_backend_mygene.py - idtrack/
_external_mappers/ , Python, 519 lines_backend_pybiomart.py - idtrack/
_external_mappers/ , Python, 283 lines_constants.py - idtrack/
_external_mappers/ , Python, 157 lines_convert.py - idtrack/
_external_mappers/ , Python, 663 lines_ortholog.py - idtrack/
_external_mappers/ , Python, 534 lines_utils.py - idtrack/
_graph_maker.py , Python, 1,160 lines - idtrack/
_harmonize_features.py , Python, 1,175 lines, 4 matches - idtrack/
_the_graph.py , Python, 1,369 lines - idtrack/
_track.py , Python, 2,300 lines - idtrack/
_track_tests.py , Python, 1,174 lines - idtrack/
_utils_hdf5.py , Python, 772 lines - idtrack/
_verify_organism.py , Python, 200 lines - noxfile.py, Python, 329 lines
- reproducibility/
experiments/ , Jupyter, 342 linesexample_manual_running.i pynb - reproducibility/
experiments/ , Jupyter, 10 linesexperiment_cellranger_id track/ Untitled.ipynb - reproducibility/
experiments/ , Jupyter, 3,325 linesexperiment_cellranger_id track/ analysis copy.ipynb - reproducibility/
experiments/ , Jupyter, 3,158 linesexperiment_cellranger_id track/ analysis.ipynb - reproducibility/
experiments/ , Jupyter, 557 linesexperiment_cellranger_id track/ analysis_gold_standard_e xpression_consistency.ip ynb - reproducibility/
experiments/ , Jupyter, 797 linesexperiment_cellranger_id track/ analysis_gold_standard_f ig4c.ipynb - reproducibility/
experiments/ , Jupyter, 405 linesexperiment_cellranger_id track/ create_data.ipynb - reproducibility/
experiments/ , Jupyter, 376 linesexperiment_external_brid ges/ 00_external_bridges_trad eoff.ipynb - reproducibility/
experiments/ , Jupyter, 1,580 linesexperiment_hlca/ 00_hlca_manuscript_table 1_and_figures.ipynb - reproducibility/
experiments/ , Jupyter, 506 linesexperiment_identifier_dr ift/ 00_identifier_drift_case _studies.ipynb - reproducibility/
experiments/ , Jupyter, 194 linesexperiment_marketing_ove rview/ 00_marketing_overview_da shboard.ipynb - reproducibility/
experiments/ , Jupyter, 279 linesexperiment_other_organis ms/ 00_other_organisms_showc ase.ipynb - reproducibility/
experiments/ , Jupyter, 351 linesexperiment_other_organis ms/ 01_other_organisms_deep_ dive_time_axis.ipynb - reproducibility/
experiments/ , Jupyter, 759 linesexperiment_random_stress _tests/ 00_random_stress_tests_f ig4ab.ipynb - reproducibility/
experiments/ , Jupyter, 291 linesexperiment_time_travel_m atrix/ 00_build_time_travel_mat rix_cache.ipynb - reproducibility/
experiments/ , Jupyter, 654 linesexperiment_time_travel_m atrix/ 01_analyze_time_travel_m atrix.ipynb - reproducibility/
experiments/ , Jupyter, 424 linesexperiment_time_travel_m atrix/ 02_roundtrip_consistency .ipynb - reproducibility/
experiments/ , Jupyter, 466 linesexperiment_time_travel_v s_external_mappers/ 00_build_time_travel_vs_ external_mappers_cache.i pynb - reproducibility/
experiments/ , Jupyter, 478 linesexperiment_time_travel_v s_external_mappers/ 01_analyze_time_travel_v s_external_mappers.ipynb - reproducibility/
experiments/ , Jupyter, 388 linesexperiment_time_travel_v s_external_mappers/ 02_case_studies_external _mapper_failures.ipynb - reproducibility/
experiments/ , Jupyter, 561 linesexperiment_tool_comparis on/ 00_tool_comparison_capab ility_matrix_fig4d.ipynb - reproducibility/
experiments/ , Jupyter, 606 linesexperiment_tool_comparis on/ 01_external_mapper_varia bility_deep_dive.ipynb - reproducibility/
experiments/ , Jupyter, 375 linesexperiment_tool_comparis on/ 02_pybiomart_release_sen sitivity.ipynb - reproducibility/
experiments/ , Jupyter, 26 linesfigure_rcparams/ rcparams.ipynb - reproducibility/
experiments/ , Jupyter, 287 linesprev/ example_manual_running.i pynb - repository limit reached (2,000 files or 30 MB): the rest is at the source (57 files)
- LICENSE, License, 29 lines
- README.rst, Text, 80 lines
cellgeni/OMERO_tools
f4d9d3f1633265743936db2cfe48b53eda992272, 28 July 2026Availability: 1 check, the latest on 28 September 2026: the link answers
- 28 September 2026: the link answers
9 files
- ann2SR.py, Python, 472 lines
- ann2Xenium.py, Python, 237 lines
- ann2Xenium_batch.py, Python, 49 lines
- ann2Xenium_coord.py, Python, 769 lines
- ann2Xenium_v2.py, Python, 184 lines
- categorical_ann2SR.py, Python, 193 lines
- duplicate.py, Python, 291 lines
- transfer_annotations_dif
ferent_OMERO_servers.py , Python, 210 lines - README.md, Text, 62 lines
haniffalab/cherita-react
740eaf25038452adf553d2528c0c8e8bcbd62906, 8 May 2026Availability: 1 check, the latest on 28 September 2026: the link answers
- 28 September 2026: the link answers
29 files
- babel.config.js, JavaScript, 18 lines
- jest.config.js, JavaScript, 21 lines
- prepare-release.js, JavaScript, 18 lines
- sites/
demo/ , JavaScript, 13 linessrc/ components/ Footer.js - sites/
demo/ , JavaScript, 16 linessrc/ index.js - sites/
demo/ , JavaScript, 40 linesvite.config.js - src/
lib/ , JavaScript, 198 linescomponents/ scatterplot/ ScatterplotLayer.js - src/
lib/ , JavaScript, 138 linesconstants/ colorscales.js - src/
lib/ , JavaScript, 111 linesconstants/ constants.js - src/
lib/ , JavaScript, 83 lineshelpers/ color-helper.js - src/
lib/ , JavaScript, 50 lineshelpers/ map-helper.js - src/
lib/ , JavaScript, 104 lineshelpers/ zarr-helper.js - src/
lib/ , JavaScript, 32 linesindex.js - src/
lib/ , JavaScript, 210 linesutils/ Filter.js - src/
lib/ , JavaScript, 294 linesutils/ Resolver.js - src/
lib/ , JavaScript, 23 linesutils/ errors.js - src/
lib/ , JavaScript, 55 linesutils/ parquetData.js - src/
lib/ , JavaScript, 104 linesutils/ requests.js - src/
lib/ , JavaScript, 67 linesutils/ search.js - src/
lib/ , JavaScript, 63 linesutils/ string.js - src/
lib/ , JavaScript, 159 linesutils/ zarrData.js - src/
lib/ , JavaScript, 7 linesviews/ ObservationFeature/ index.js - src/
lib/ , JavaScript, 7 linesviews/ PerturbationMap/ index.js - src/
lib/ , JavaScript, 45 linesworkers/ scatterplotData.js - website/
docusaurus.config.js , JavaScript, 116 lines - website/
sidebars.js , JavaScript, 35 lines - website/
src/ , JavaScript, 5 linespages/ index.js - LICENSE, License, 21 lines
- README.md, Text, 36 lines
theislab/sctram
e570de5f9ec6dc783542b82c4f88b192677d4347, 10 April 2026Availability: 1 check, the latest on 28 September 2026: the link answers
- 28 September 2026: the link answers
35 files
- docs/
_notebooks/ , Jupyter, 13 lines_doctest.ipynb - docs/
conf.py , Python, 181 lines - noxfile.py, Python, 235 lines
- reproducibility/
figure_rcparams/ , Jupyter, 23 linesrcparams.ipynb - reproducibility/
local/ , Jupyter, 1,813 lines, 6 matchesexperiments/ sctram_figures_only.ipyn b - reproducibility/
modules/ , Jupyter, 90 linesmodule_api/ example_operations.ipynb - reproducibility/
modules/ , Jupyter, 152 lines, 1 matchmodule_evaluate/ example_operations_adjac ency.ipynb - reproducibility/
modules/ , Jupyter, 233 linesmodule_evaluate/ example_operations_conve rters.ipynb - reproducibility/
modules/ , Jupyter, 134 linesmodule_evaluate/ example_operations_embed ding.ipynb - reproducibility/
modules/ , Jupyter, 222 linesmodule_evaluate/ example_operations_pseud otimevalues.ipynb - reproducibility/
modules/ , Jupyter, 134 linesmodule_evaluate/ example_operations_spati al_embedding.ipynb - reproducibility/
modules/ , Jupyter, 160 linesmodule_evaluate/ example_operations_spati al_pseudotime.ipynb - reproducibility/
modules/ , Jupyter, 83 linesmodule_generate/ submodule_real/ example_operations.ipynb - reproducibility/
modules/ , Jupyter, 400 lines, 2 matchesmodule_generate/ submodule_synthetic_simp le/ example_operations.ipynb - reproducibility/
modules/ , Jupyter, 96 linesmodule_infer/ example_operations_embed ding.ipynb - reproducibility/
modules/ , Jupyter, 224 linesmodule_infer/ example_operations_infer ence.ipynb - reproducibility/
modules/ , Jupyter, 78 linesmodule_input/ bootstapping_plotting.ip ynb - reproducibility/
modules/ , Jupyter, 437 linesmodule_input/ example_operations.ipynb - reproducibility/
modules/ , Jupyter, 86 linesput_somewhere.ipynb - reproducibility/
preprocessing_real_data/ , Jupyter, 283 linesnorman_sciplex_cpa.ipynb - reproducibility/
preprocessing_real_data/ , Jupyter, 57 linesother_developmental_comp letal.ipynb - reproducibility/
preprocessing_real_data/ , Jupyter, 112 linessuo_developmental_comple te.ipynb - reproducibility/
scripts/ , Python, 59 linescheck_occurances_for_ver sion.py - reproducibility/
scripts/ , Shell, 68 linesjupyter_tunnel.sh - reproducibility/
scripts/ , Shell, 270 linessync_server.sh - reproducibility/
scripts/ , Shell, 6 linesutils/ _colors.sh - reproducibility/
scripts/ , Shell, 5 linesutils/ _log.sh - reproducibility/
scripts/ , Shell, 31 linesutils/ _python_environment.sh - reproducibility/
scripts/ , Python, 322 linesutils/ _system.py - reproducibility/
server_sync/ , Python, 101 lines, 2 matchesexperiments/ _deprecated_suo_incremen tal/ helper/ model_training.py - reproducibility/
server_sync/ , Python, 137 lines, 1 matchexperiments/ _deprecated_suo_incremen tal/ helper/ suo_scib_iterative.py - reproducibility/
server_sync/ , Python, 89 linesexperiments/ _deprecated_suo_incremen tal/ helper/ suo_sctram_iterative.py - reproducibility/
server_sync/ , Jupyter, 502 lines, 3 matchesexperiments/ _deprecated_suo_incremen tal/ incremental_analysis.ipy nb - repository limit reached (2,000 files or 30 MB): the rest is at the source (151 files)
- LICENSE, License, 29 lines
- README.rst, Text, 55 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 5 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 129 scripts, each with its path and the digest of its content;
- 19 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data and code availability
Data: Raw de novo whole embryo data will be available on EMBL-EBI ArrayExpress and BioImage Archive. All data accessions are currently private, and will be made available upon publication:
scRNAseq raw FASTQ and counts (embryo F137), E-MTAB-11504
scRNAseq raw FASTQ and counts (embryo F147), E-MTAB-11520
scRNAseq raw FASTQ and counts (embryo F158), E-MTAB-11911
Multiome raw FASTQ and counts (embryos F137, F147, F158), E-MTAB-16757
Spatial transcriptomics Visium raw FASTQs, counts, images (embryo SOB26), E-MTAB-14785
Spatial transcriptomics Xenium raw cell x gene counts post Xenium cell segmentation, images (embryos SOB25, SOB26), BioImage Archive, S-BIAD2878
The HDCA atlas will be accessible via an interactive web portal at https://
Code: All code will be made publicly available at https://
Additional repositories associated with this study include https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 28 September 2026: the first record
Recorded: type, journal, dates, 94 authors, 5 funders, 189 references.
Cite
This paper
Webb, S., Rose, A., Xu, C., Steele, L., Kuri, M. A., Stephenson, E., Inecik, K., Jafree, D., Foster, A. R., Basurto-Lozada, D., Chipampe, N.-J., Pournara, A. V., Jacques, M.-A., To, K., Admane, C., Kritikaki, E., Chroscik, M. A., Horsfall, D., Foreman, J., . . . Haniffa, M. (2026). An integrated single cell and spatial omics atlas of human prenatal development. bioRxiv (preprint). https://
BibTeX
@article{webb2026integra
author = {Webb, Simone and Rose, Antony and Xu, Chuan and Steele, Lloyd and Kuri, Masami A and Stephenson, Emily and Inecik, Kemal and Jafree, Daniyal and Foster, April R and Basurto-Lozada, Daniela and Chipampe, Nana-Jane and Pournara, Anna V and Jacques, Marc-Antoine and To, Ken and Admane, Chloe and Kritikaki, Efpraxia and Chroscik, Marta A and Horsfall, David and Foreman, Julia and Rademaker, Koen and Karjalainen, Julia and Laddach, Anna and Madad, Shaista and Lawrence, John EG and Kleshchevnikov, Vitalii and Lisgo, Steven and Lee, Jimmy Tsz Hang and Blevinal, Joséphine and Alqahtani, Ahlam and Makarchuk, Stanislaw and Jackson, Jake and Ucuncu, Ekin and Silva, Teresa P and Lorenzi, Valentina and Torabi, Fereshteh and Botting, Rachel A and Roberts, Kenny and Olabi, Bayanne and Chakala, Keerthi P and Dony, Leander and Dall’Aglio, Gabriele and Cujba, Ana-Maria and Whitfield, Holly J and Hale, Christine and Engelbert, Justin and Li, Tong and Prigmore, Elena and Gopee, Nusayhah Hudaa and Mather, Michael W and Talks, Ben and Fei, Lijiang and Pereira, Diana and Moldovan, Corina and Elmentaite, Rasa and Polanski, Krzysztof and Predeus, Alexander V and Mazin, Pasha and Rowe, Vicky and Bayraktar, Semih and Knight-Schrijver, Vincent R and Clatworthy, Menna and Kanemaru, Kazumasa and Cranley, James and Guo, Andrew and Sinha, Sanjay and McDonald, David and Filby, Andrew and Kerwin, Janet and He, Peng and Sequeira, Inês and Vento-Tormo, Roser and Luecken, Malte D. and Chédotal, Alain and Taylor, Deanne and Martin-Almedina, Silvia and Ostergaard, Pia and Behjati, Sam and Robertson, Elizabeth and Wilson, Nicola K and Marioni, John and Firth, Helen V and Wright, Caroline F and Pourquie, Olivier and Gambardella, Laure and Alexandre, Paula and Gottgens, Berthold and Meyer, Kerstin B and Vijayabaskar, MS and Lotfollahi, Mo and Teichmann, Sarah A and Theis, Fabian J and Progatzky, Fränze and Bayraktar, Omer Ali and Haniffa, Muzlifah},
title = {{An integrated single cell and spatial omics atlas of human prenatal development}},
journal = {bioRxiv (preprint)},
year = {2026},
month = apr,
publisher = {bioRxiv},
issn = {2692-8205},
doi = {10.64898/
url = {https://
}
RIS
TY - JOUR
AU - Webb, Simone
AU - Rose, Antony
AU - Xu, Chuan
AU - Steele, Lloyd
AU - Kuri, Masami A
AU - Stephenson, Emily
AU - Inecik, Kemal
AU - Jafree, Daniyal
AU - Foster, April R
AU - Basurto-Lozada, Daniela
AU - Chipampe, Nana-Jane
AU - Pournara, Anna V
AU - Jacques, Marc-Antoine
AU - To, Ken
AU - Admane, Chloe
AU - Kritikaki, Efpraxia
AU - Chroscik, Marta A
AU - Horsfall, David
AU - Foreman, Julia
AU - Rademaker, Koen
AU - Karjalainen, Julia
AU - Laddach, Anna
AU - Madad, Shaista
AU - Lawrence, John EG
AU - Kleshchevnikov, Vitalii
AU - Lisgo, Steven
AU - Lee, Jimmy Tsz Hang
AU - Blevinal, Joséphine
AU - Alqahtani, Ahlam
AU - Makarchuk, Stanislaw
AU - Jackson, Jake
AU - Ucuncu, Ekin
AU - Silva, Teresa P
AU - Lorenzi, Valentina
AU - Torabi, Fereshteh
AU - Botting, Rachel A
AU - Roberts, Kenny
AU - Olabi, Bayanne
AU - Chakala, Keerthi P
AU - Dony, Leander
AU - Dall’Aglio, Gabriele
AU - Cujba, Ana-Maria
AU - Whitfield, Holly J
AU - Hale, Christine
AU - Engelbert, Justin
AU - Li, Tong
AU - Prigmore, Elena
AU - Gopee, Nusayhah Hudaa
AU - Mather, Michael W
AU - Talks, Ben
AU - Fei, Lijiang
AU - Pereira, Diana
AU - Moldovan, Corina
AU - Elmentaite, Rasa
AU - Polanski, Krzysztof
AU - Predeus, Alexander V
AU - Mazin, Pasha
AU - Rowe, Vicky
AU - Bayraktar, Semih
AU - Knight-Schrijver, Vincent R
AU - Clatworthy, Menna
AU - Kanemaru, Kazumasa
AU - Cranley, James
AU - Guo, Andrew
AU - Sinha, Sanjay
AU - McDonald, David
AU - Filby, Andrew
AU - Kerwin, Janet
AU - He, Peng
AU - Sequeira, Inês
AU - Vento-Tormo, Roser
AU - Luecken, Malte D.
AU - Chédotal, Alain
AU - Taylor, Deanne
AU - Martin-Almedina, Silvia
AU - Ostergaard, Pia
AU - Behjati, Sam
AU - Robertson, Elizabeth
AU - Wilson, Nicola K
AU - Marioni, John
AU - Firth, Helen V
AU - Wright, Caroline F
AU - Pourquie, Olivier
AU - Gambardella, Laure
AU - Alexandre, Paula
AU - Gottgens, Berthold
AU - Meyer, Kerstin B
AU - Vijayabaskar, MS
AU - Lotfollahi, Mo
AU - Teichmann, Sarah A
AU - Theis, Fabian J
AU - Progatzky, Fränze
AU - Bayraktar, Omer Ali
AU - Haniffa, Muzlifah
TI - An integrated single cell and spatial omics atlas of human prenatal development
T2 - bioRxiv (preprint)
J2 - bioRxiv
PY - 2026
DA - 2026/
SN - 2692-8205
PB - bioRxiv
DO - 10.64898/
UR - https://
ER -
CSL-JSON
{
"id": "10.64898/
"type": "article",
"title": "An integrated single cell and spatial omics atlas of human prenatal development",
"container-title": "bioRxiv (preprint)",
"author": [
{
"family": "Webb",
"given": "Simone"
},
{
"family": "Rose",
"given": "Antony"
},
{
"family": "Xu",
"given": "Chuan"
},
{
"family": "Steele",
"given": "Lloyd"
},
{
"family": "Kuri",
"given": "Masami A"
},
{
"family": "Stephenson",
"given": "Emily"
},
{
"family": "Inecik",
"given": "Kemal"
},
{
"family": "Jafree",
"given": "Daniyal"
},
{
"family": "Foster",
"given": "April R"
},
{
"family": "Basurto-Lozada",
"given": "Daniela"
},
{
"family": "Chipampe",
"given": "Nana-Jane"
},
{
"family": "Pournara",
"given": "Anna V"
},
{
"family": "Jacques",
"given": "Marc-Antoine"
},
{
"family": "To",
"given": "Ken"
},
{
"family": "Admane",
"given": "Chloe"
},
{
"family": "Kritikaki",
"given": "Efpraxia"
},
{
"family": "Chroscik",
"given": "Marta A"
},
{
"family": "Horsfall",
"given": "David"
},
{
"family": "Foreman",
"given": "Julia"
},
{
"family": "Rademaker",
"given": "Koen"
},
{
"family": "Karjalainen",
"given": "Julia"
},
{
"family": "Laddach",
"given": "Anna"
},
{
"family": "Madad",
"given": "Shaista"
},
{
"family": "Lawrence",
"given": "John EG"
},
{
"family": "Kleshchevnikov",
"given": "Vitalii"
},
{
"family": "Lisgo",
"given": "Steven"
},
{
"family": "Lee",
"given": "Jimmy Tsz Hang"
},
{
"family": "Blevinal",
"given": "Joséphine"
},
{
"family": "Alqahtani",
"given": "Ahlam"
},
{
"family": "Makarchuk",
"given": "Stanislaw"
},
{
"family": "Jackson",
"given": "Jake"
},
{
"family": "Ucuncu",
"given": "Ekin"
},
{
"family": "Silva",
"given": "Teresa P"
},
{
"family": "Lorenzi",
"given": "Valentina"
},
{
"family": "Torabi",
"given": "Fereshteh"
},
{
"family": "Botting",
"given": "Rachel A"
},
{
"family": "Roberts",
"given": "Kenny"
},
{
"family": "Olabi",
"given": "Bayanne"
},
{
"family": "Chakala",
"given": "Keerthi P"
},
{
"family": "Dony",
"given": "Leander"
},
{
"family": "Dall’Aglio",
"given": "Gabriele"
},
{
"family": "Cujba",
"given": "Ana-Maria"
},
{
"family": "Whitfield",
"given": "Holly J"
},
{
"family": "Hale",
"given": "Christine"
},
{
"family": "Engelbert",
"given": "Justin"
},
{
"family": "Li",
"given": "Tong"
},
{
"family": "Prigmore",
"given": "Elena"
},
{
"family": "Gopee",
"given": "Nusayhah Hudaa"
},
{
"family": "Mather",
"given": "Michael W"
},
{
"family": "Talks",
"given": "Ben"
},
{
"family": "Fei",
"given": "Lijiang"
},
{
"family": "Pereira",
"given": "Diana"
},
{
"family": "Moldovan",
"given": "Corina"
},
{
"family": "Elmentaite",
"given": "Rasa"
},
{
"family": "Polanski",
"given": "Krzysztof"
},
{
"family": "Predeus",
"given": "Alexander V"
},
{
"family": "Mazin",
"given": "Pasha"
},
{
"family": "Rowe",
"given": "Vicky"
},
{
"family": "Bayraktar",
"given": "Semih"
},
{
"family": "Knight-Schrijver",
"given": "Vincent R"
},
{
"family": "Clatworthy",
"given": "Menna"
},
{
"family": "Kanemaru",
"given": "Kazumasa"
},
{
"family": "Cranley",
"given": "James"
},
{
"family": "Guo",
"given": "Andrew"
},
{
"family": "Sinha",
"given": "Sanjay"
},
{
"family": "McDonald",
"given": "David"
},
{
"family": "Filby",
"given": "Andrew"
},
{
"family": "Kerwin",
"given": "Janet"
},
{
"family": "He",
"given": "Peng"
},
{
"family": "Sequeira",
"given": "Inês"
},
{
"family": "Vento-Tormo",
"given": "Roser"
},
{
"family": "Luecken",
"given": "Malte D."
},
{
"family": "Chédotal",
"given": "Alain"
},
{
"family": "Taylor",
"given": "Deanne"
},
{
"family": "Martin-Almedina",
"given": "Silvia"
},
{
"family": "Ostergaard",
"given": "Pia"
},
{
"family": "Behjati",
"given": "Sam"
},
{
"family": "Robertson",
"given": "Elizabeth"
},
{
"family": "Wilson",
"given": "Nicola K"
},
{
"family": "Marioni",
"given": "John"
},
{
"family": "Firth",
"given": "Helen V"
},
{
"family": "Wright",
"given": "Caroline F"
},
{
"family": "Pourquie",
"given": "Olivier"
},
{
"family": "Gambardella",
"given": "Laure"
},
{
"family": "Alexandre",
"given": "Paula"
},
{
"family": "Gottgens",
"given": "Berthold"
},
{
"family": "Meyer",
"given": "Kerstin B"
},
{
"family": "Vijayabaskar",
"given": "MS"
},
{
"family": "Lotfollahi",
"given": "Mo"
},
{
"family": "Teichmann",
"given": "Sarah A"
},
{
"family": "Theis",
"given": "Fabian J"
},
{
"family": "Progatzky",
"given": "Fränze"
},
{
"family": "Bayraktar",
"given": "Omer Ali"
},
{
"family": "Haniffa",
"given": "Muzlifah"
}
],
"container-title-short":
"DOI": "10.64898/
"ISSN": "2692-8205",
"publisher": "bioRxiv",
"URL": "https://
"issued": {
"date-parts": [
[
2026,
4,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: Squidpy, CuPy, UMAP, 12 other tools, 3 references
- [2] doi:10.1038/s41467-026-68596-w [code]
- Spatial cartography of human thymus enables the geopositioning of lineage transcription factors in rare mimetic thymic epithelial cells.Journal: Nature communicationsIn common: Squidpy, anndata, Scanpy, 9 other tools, 5 references
- [3] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: CuPy, Biopython, UMAP, 12 other tools, 2 references
- [4] doi:10.1016/j.celrep.2026.117270 [code]
- Blocking apoptosis promotes survival and alters developmental dynamics of human retinal ganglion cells in retinal organoids.Journal: Cell reportsIn common: UMAP, anndata, Scanpy, 7 other tools, 6 references
- [5] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: scVelo, UMAP, anndata, 10 other tools, 3 references
- [6] doi:10.1016/j.crmeth.2026.101342 [code]
- Interpretable learning of temporal cellular dynamics from single-cell data.Journal: Cell reports methodsIn common: scVelo, UMAP, anndata, 11 other tools, 1 reference
- [7] doi:10.21203/rs.3.rs-9676637/v1 [code]
- A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic TechnologiesJournal: Research Square (preprint)In common: Squidpy, UMAP, anndata, 10 other tools, 2 references
- [8] doi:10.1038/s44320-026-00208-7 [code]
- Interpretable deep generative ensemble learning for single-cell omics with Hydra.Journal: Molecular systems biologyIn common: UMAP, anndata, Scanpy, 8 other tools, 5 references
- [9] doi:10.1016/j.xgen.2026.101217 [code]
- ProtoCloud: A prototypical self-explaining model for single-cell analysis.Journal: Cell genomicsIn common: UMAP, anndata, Scanpy, 8 other tools, 5 references
- [10] doi:10.1093/nar/gkag706 [code]
- scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.Journal: Nucleic acids researchIn common: Squidpy, UMAP, anndata, 11 other tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 5 repositories of the authors' code, each at its verified commit and with its license, 129 scripts, and 19 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:f240b2a90a0a1070…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
