OSCR

Brachyury expression levels predict lineage potential and axis-forming ability of in vitro-derived neuromesodermal progenitors.

Code ↔ Paper

5 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 5 matches
  1. [1] § Materials and methods › Spatial gene expression pattern analysis for in vivo data ↔ Figure 6E.ipynb, lines 51–206 · score 0.73 · covariance matrix, expression profiles, bootstrap, gi, discretised, python
  2. [2] § Results › 6. TBXT and SOX2 in vivo spatial patterns provide separate information to NM fated regions ↔ Figure 6E.ipynb, lines 447–493 · score 0.64 · cell width, positional error, AP midline, AP axis, genes, ratio
  3. [3] § Materials and methods › Mouse husbandry ↔ data-analysis-math-model.ipynb, lines 1–45 · score 0.61 · Biosystems Dynamics Research, RIKEN Center, Edinburgh, Laboratory, UK, Animal
  4. [4] § Materials and methods › Gastruloid culture, live imaging and quantitative analysis ↔ Figure 4B-E.ipynb, lines 463–584 · score 0.61 · gastruloid trace, NMP region, DMSO, isolating, median, orientation
  5. [5] § Materials and methods › Gastruloid culture, live imaging and quantitative analysis ↔ batch_scripts/NMP_ROI_v3.py, lines 280–369 · score 0.55 · median projecting, poles, centroid, isolating, stacks, orientation

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 654 lines · 19 KB · MIT · 2 matches

  1. # %%
  2. import numpy as np
  3. import matplotlib.pyplot as plt
  4. import pandas as pd
  5. import os
  6. import seaborn as sns
  7. import matplotlib.pyplot as plt
  8. from scipy.ndimage import uniform_filter1d
  9. from scipy.ndimage import gaussian_filter1d
  10. import warnings
  11. from scipy.spatial import cKDTree
  12. warnings.filterwarnings('ignore')
  13. # Import Wt_E8_data data from French et al 2025 data file 1
  14. Wt_E8_fileName = 'Data_1.xlsx'
  15. Wt_E8_sheetName = 'Wt_E8.5_data'
  16. Wt_E8_data = pd.read_excel(
  17. Wt_E8_fileName,
  18. sheet_name=Wt_E8_sheetName
  19. )
  20. # Also import data that contains real distances AP disitance from notochord
  21. distances = pd.read('Data_4.csv')
  22. # merge datasets
  23. all_cells = pd.concat([all_cells, distances], axis=1)
  24. # Normalise data to 0-1 range - primarily for plotting purposes
  25. # and putting all genes on the same scale.
  26. # Calculations not be sensitive to unit values by design.
  27. tbxt_percentiles = np.percentile(all_cells['TBXT'], (0.1, 99.99))
  28. sox2_percentiles = np.percentile(all_cells['SOX2'], (0.1, 99.99))
  29. all_cells['TBXT_01_norm'] = (all_cells['TBXT'] - tbxt_percentiles[0]) / (tbxt_percentiles[1] - tbxt_percentiles[0])
  30. all_cells['SOX2_01_norm'] = (all_cells['SOX2'] - sox2_percentiles[0]) / (sox2_percentiles[1] - sox2_percentiles[0])
  31. tbxt_percentiles = np.percentile(all_cells['log_TBXT'], (0.1, 99.99))
  32. sox2_percentiles = np.percentile(all_cells['log_SOX2'], (0.1, 99.99))
  33. all_cells['TBXT_01_norm_log'] = (all_cells['log_TBXT'] - tbxt_percentiles[0]) / (tbxt_percentiles[1] - tbxt_percentiles[0])
  34. all_cells['SOX2_01_norm_log'] = (all_cells['log_SOX2'] - sox2_percentiles[0]) / (sox2_percentiles[1] - sox2_percentiles[0])
  35. # Define ratios from differenet calculations
  36. all_cells['TBXT_SOX2_ratio'] = np.log2((all_cells['TBXT']/all_cells['SOX2']))
  37. all_cells['TBXT_SOX2_ratio_log'] = all_cells['SOX2_01_norm_log'] - all_cells['TBXT_01_norm_log']
  38. # %% [markdown]
  39. # ### Original Dubuis et al 2013 matlab function
  40. #
  41. # %% Computes positional error carried by nG genes. Assumes that at every position, the
  42. # %% gene expression variability is well described by a multivariate gaussian. Bootstraps
  43. # %% the result be repeatedly estimating the error with leave-one-out procedure.
  44. # %
  45. # %
  46. # % Julien Dubuis, Gasper Tkacik (2008, 2010, 2014)
  47. # %
  48. # % INPUT:
  49. # % data
  50. # % nG x nN x nX matrix containing samples of expression profiles
  51. # % for nG genes from nN embroys, sampled at positions nX
  52. # % data should be normalized such that the mean profile is between
  53. # % 0 and 1 (individual profiles can fluctuate out of these bounds)
  54. # % nbins
  55. # % the new discretization to use along the position axis; nbins
  56. # % has to divide nX without a reminder; consecutive bins in the
  57. # % original discretization will be averaged together
  58. # % zsmooth
  59. # % smoothing to apply to the mean profiles before computing the
  60. # % numerical derivative; indicates the number of consecutive bins
  61. # % to smooth over; 1 = no smoothing
  62. # % OUTPUT:
  63. # % sigmaX
  64. # % estimate of positional error, of dimension (nshuffles) x
  65. # % (nbins), containing nshuffles leave-one-out estimates
  66. # % meanG
  67. # % the mean of the profiles using the new smoothing and
  68. # % discretization; dimension nG x nbins
  69. # % dGdX
  70. # % numerical spatial derivative of the gene expression patterns;
  71. # % dimension nG x nbins
  72. # % covG
  73. # % the covariance matrix of expression levels of dimension nG x nG
  74. # % x nbins
  75. #
  76. # function [sigmaX meanG dGdX covG] = compute_SigmaX(data, nbins, zsmooth)
  77. #
  78. # if (nargin <= 1)
  79. # nbins = 100;
  80. # zsmooth = 1;
  81. # end;
  82. #
  83. # nshuffles = 10;
  84. #
  85. # if (numel(size(data)) == 2)
  86. # data = reshape(data, 1, size(data,1), size(data, 2));
  87. # end
  88. #
  89. # [nG nN nX] = size(data);
  90. #
  91. # if (mod(nX, nbins)~=0)
  92. # error('The new number of spatial bins must divide the original number.');
  93. # end
  94. #
  95. # disp(sprintf('**** COMPUTE_SIGMAX: Estimating positional error of %d genes from %d embryos.', nG, nN));
  96. # disp(sprintf('**** COMPUTE_SIGMAX: %d spatial bins -> %d spatial bins, smoothing over %d bins.', nX, nbins, zsmooth));
  97. #
  98. #
  99. # nK = nX/nbins;
  100. #
  101. # dX = 1./nbins;
  102. #
  103. # meanG = zeros(nG, nshuffles, nbins);
  104. #
  105. # covgg=cell(nshuffles, nbins);
  106. #
  107. # divgg=cell(nshuffles, nbins);
  108. #
  109. # sigmax4G=NaN(nshuffles, nbins);
  110. #
  111. # for kk=1:nshuffles
  112. #
  113. # samps=randperm(nN);
  114. # samps=samps(1:nN-1);
  115. #
  116. # gg = data(:,samps,:);
  117. #
  118. # for l=1:nbins,
  119. #
  120. # meanG(:,kk,l) = squeeze(nanmean(nanmean(gg(:,:,(l-1)*nK+1:l*nK),2),3));
  121. #
  122. # for q=1:nG,
  123. # tt(q,:) = flatten(gg(q,:,(l-1)*nK+1:l*nK));
  124. # end
  125. #
  126. # covgg{kk,l}=nancov(tt');
  127. #
  128. # end
  129. #
  130. # if (~isempty(zsmooth))
  131. # for q=1:nG,
  132. # meanG(q,kk,:) = smooth(meanG(q,kk,:),zsmooth);
  133. # end
  134. # end
  135. #
  136. # divgg{kk,1} = (meanG(:,kk,2) - meanG(:,kk,1))/dX;
  137. # divgg{kk,nbins} = (meanG(:,kk,nbins) - meanG(:,kk,nbins-1))/dX;
  138. #
  139. # for l=2:nbins-1,
  140. # divgg{kk,l} = (meanG(:,kk,l+1)-meanG(:,kk,l-1))/(2*dX);
  141. # end
  142. #
  143. #
  144. # %size(divgg{kk,l})
  145. # for l=1:nbins,
  146. # if (sum(flatten(isnan(covgg{kk,l}))) > 0 || sum(isnan(divgg{kk,l}))>0)
  147. # sigmaX(kk,l) = nan;
  148. # else
  149. # sigmaX(kk,l) = 1/sqrt(divgg{kk,l}'*covgg{kk,l}^(-1)*divgg{kk,l});
  150. # end
  151. # end
  152. # end
  153. #
  154. # for l=1:nbins,
  155. # covG(:,:,l) = covgg{1,l};
  156. # dGdX(:,l) = divgg{1,l};
  157. # end
  158. # meanG = squeeze(mean(meanG(:,:,:),2));
  159. # end
  160. # %% [markdown]
  161. # ### Python translation
  162. # %%
  163. def make_data_array(df, cols, bin_col="LR_bin"):
  164. """"
  165. Convert dataframe to 3D array of shape (nGenes, nEmbryos, nPositions) for positional error calculation.
  166. """
  167. embryos = np.sort(df["Embryo"].unique()) # sort embryos to ensure consistent ordering
  168. positions = np.sort(df[bin_col].unique()) # sort positions to ensure consistent ordering
  169. # Initialize array with NaN values to handle missing data
  170. data = np.full(
  171. (len(cols), len(embryos), len(positions)),
  172. np.nan
  173. )
  174. # Create lookup dictionaries to map embryo and position labels to array indices
  175. embryo_lookup = {e:i for i,e in enumerate(embryos)}
  176. pos_lookup = {p:i for i,p in enumerate(positions)}
  177. for _, row in df.iterrows():
  178. ei = embryo_lookup[row["Embryo"]]
  179. pi = pos_lookup[row[bin_col]]
  180. for gi, col in enumerate(cols):
  181. data[gi, ei, pi] = row[col]
  182. return data, positions, embryos
  183. def positional_error(data, smoothing=None):
  184. """
  185. Calculate positional error using the method of Dubuis et al 2013,
  186. Direct tranlsation of MATLAB code from Dubuis et al 2013.
  187. Minor changes to pre-compute calculations outside of loops.
  188. """
  189. # data array is nGenes x nEmbryos x nX-positions
  190. nG, nE, nX = data.shape
  191. # No. shuffles for leave-one-out procedure
  192. # Same as in Dubuis et al 2013
  193. nshuffles = 10
  194. # Initialize array to hold positional error estimates for each shuffle and position
  195. sigma = np.full((nshuffles, nX), np.nan)
  196. # Loop over shuffles
  197. for kk in range(nshuffles):
  198. # Randomly permute embryos and leave one out
  199. perm = np.random.permutation(nE)
  200. samps = perm[:nE - 1]
  201. gg = data[:, samps, :]
  202. # Mean profile for this shuffle
  203. # this is calculated in loop in original
  204. # MATLAB code but can be pre-computed here
  205. meanG_kk = np.nanmean(gg, axis=1) # nG x nX
  206. # Smooth mean profile if requested
  207. # Typically smoothing is set to 3 as in Dubuis et al 2013
  208. # I use smaller gaussian smoothing to remove high
  209. # frequency noise rather than wider uniform smoothing as in Dubuis et al 2013.
  210. # Their image collection method introduces some smoothing anyway
  211. # than single cell resoltion we have here.
  212. if smoothing is not None and smoothing > 1:
  213. for g in range(nG):
  214. meanG_kk[g, :] = uniform_filter1d(
  215. meanG_kk[g, :],
  216. size=smoothing,
  217. mode='nearest'
  218. )
  219. # Pre-compute gradient of mean profile with respect to position
  220. # This is the "g'" term in the positional error formula and
  221. # calculate in loop in the original MATLAB code
  222. dGdX = np.gradient(meanG_kk, axis=1)
  223. # Loop over positions to calculate positional error at each position
  224. for x in range(nX):
  225. # Extract data for this position across all genes and embryos
  226. X = data[:, :, x] # shape (nG, nE)
  227. # Remove any genes with NaN values at this position
  228. valid = ~np.any(np.isnan(X), axis=0)
  229. X = X[:, valid]
  230. # Find covariance matrix of gene expression at this position
  231. covgg = np.cov(X)
  232. if nG == 1:
  233. covgg = np.array([[covgg]])
  234. try:
  235. # Calculate error using formula from Dubuis et al 2013:
  236. # error = g' * C^(-1) * g
  237. error = (
  238. dGdX[:, x].T # use .T to ensure g is a row vector
  239. @ np.linalg.inv(covgg) # C^(-1) - use linear algebra inversion to get inverse of covariance matrix
  240. @ dGdX[:, x] # g
  241. )
  242. # Invert error to get positional error or bin-bin uncertainty estimate (sigma)
  243. # sigma = 1 / sqrt( g' * C^(-1) * g )
  244. sigma[kk,x] = 1 / np.sqrt(error)
  245. except np.linalg.LinAlgError:
  246. pass
  247. return sigma
  248. # %%
  249. # Isolate cells in somite stage 4 and 6
  250. # These are stages where we have fate data
  251. epi_nuc = all_cells[(all_cells['SP'] == 4) | (all_cells['SP'] == 6)]
  252. # Set bin size to approximate width of epithelia cell
  253. distances_list = []
  254. for embryo in epi_nuc['Embryo'].unique():
  255. embryo_data = epi_nuc[epi_nuc['Embryo'] == embryo].copy()
  256. tree = cKDTree(embryo_data[['X', 'Y', 'Z']])
  257. distances, _ = tree.query(embryo_data[['X', 'Y', 'Z']], k=6)
  258. mean_distance = np.mean(distances[:, 1]) # mean distance to 5 nearest neighbors (excluding self)
  259. distances_list.append(mean_distance)
  260. # Calculate average and standard deviation of distances to nearest neighbors across all embryos
  261. distances_list = np.array(distances_list)
  262. ap_bin_size = np.mean(distances_list)
  263. std_bin_size = np.std(distances_list)
  264. print(f"Average cell width (bin size) = {ap_bin_size:.2f} microns, std = {std_bin_size:.2f} microns")
  265. # Isolate midline cells - witthin regions of interest
  266. epi_nuc_ap_midline = epi_nuc[(np.abs(epi_nuc['Dist_to_Midline']) <= 25)]
  267. epi_nuc_medial_lateral = epi_nuc[((epi_nuc['Dist_to_Notochord']) >= 0) & ((epi_nuc['Dist_to_Notochord']) <= 50)]
  268. genesets = {
  269. "SOX2": ["SOX2_01_norm"],
  270. "TBXT": ["TBXT_01_norm"],
  271. "SOX2_TBXT": ["SOX2_01_norm", "TBXT_01_norm"],
  272. "RATIO": ["TBXT_SOX2_ratio_log"],
  273. }
  274. # cut the AP axis into bins of size ap_bin_size and find the average expression of each channel in each bin, then plot the average expression against the midline AP position for each channel
  275. epi_nuc_ap_midline['AP_bin'] = (epi_nuc_ap_midline['Dist_to_Notochord'] // ap_bin_size) * ap_bin_size
  276. epi_nuc_medial_lateral['LR_bin'] = (epi_nuc_medial_lateral['Dist_to_Midline'] // ap_bin_size) * ap_bin_size
  277. # subset between -100 and 100 on the AP axis
  278. epi_nuc_medial_lateral = epi_nuc_medial_lateral[(epi_nuc_medial_lateral['Dist_to_Notochord'] >= -150) & (epi_nuc_medial_lateral['Dist_to_Notochord'] <= 150)]
  279. epi_nuc_medial_lateral = epi_nuc_medial_lateral[(epi_nuc_medial_lateral['Dist_to_Midline'] >= -180) & (epi_nuc_medial_lateral['Dist_to_Midline'] <= 180)]
  280. # group by AP bin and find average expression of each channel in each bin
  281. epi_nuc_ap_midline_avg = epi_nuc_ap_midline.groupby(['AP_bin', 'Embryo']).agg({
  282. 'SOX2_01_norm': 'mean',
  283. 'TBXT_01_norm': 'mean',
  284. 'TBXT_SOX2_ratio_log': 'mean'
  285. }).reset_index()
  286. epi_nuc_medial_lateral_avg = epi_nuc_medial_lateral.groupby(['LR_bin', 'Embryo']).agg({
  287. 'SOX2_01_norm': 'mean',
  288. 'TBXT_01_norm': 'mean',
  289. 'TBXT_SOX2_ratio_log': 'mean'
  290. }).reset_index()
  291. # Kernel width of smoothing used in Dubuis et al 2013 is 3 cells,
  292. smoothing_internal = 3
  293. # =========== First calculate positional error along AP axis for midline cells ===========
  294. results_AP_midline = {}
  295. for name, cols in genesets.items():
  296. data, positions, embryos = make_data_array(
  297. epi_nuc_ap_midline_avg,
  298. cols,
  299. bin_col="AP_bin"
  300. )
  301. sigma= (
  302. positional_error(
  303. data,
  304. smoothing=smoothing_internal
  305. )
  306. )
  307. # calculate SEM as std/sqrt(n) where n is the number of shuffles (nshuffles)
  308. nshuffles = sigma.shape[0]
  309. mean_sigma = np.nanmean(sigma, axis=0)
  310. sem_sigma = np.nanstd(sigma, axis=0) / np.sqrt(nshuffles)
  311. lower = mean_sigma - 1.96 * sem_sigma
  312. upper = mean_sigma + 1.96 * sem_sigma
  313. results_AP_midline[name] = {
  314. "AP_bin": positions,
  315. "sigma": mean_sigma,
  316. "lower": lower,
  317. "upper": upper
  318. }
  319. notostart =0 # position of notochord on AP axis - set to 0 as we are plotting relative to notochord
  320. ROIdepth = 50 # apprximate width of region - about the width of notochord
  321. fig, ax = plt.subplots(2,2, figsize=(3, 2), sharex=True)
  322. ax = ax.flatten()
  323. colors = {
  324. "SOX2": "#D65BC6",
  325. "TBXT": "#28b834",
  326. "SOX2_TBXT": "#000000",
  327. "RATIO": "#db4444"
  328. }
  329. # Add shaded regions to indicate notochord and paraxial mesoderm - these are the same in all plots so can be added in loop
  330. for axs in ax:
  331. axs.axvspan(
  332. notostart,
  333. notostart + ROIdepth,
  334. color='#5ab4ac',
  335. alpha=0.5
  336. )
  337. axs.axvspan(
  338. notostart,
  339. notostart - ROIdepth,
  340. color="#808080",
  341. alpha=0.5
  342. )
  343. axs.axvspan(
  344. notostart- ROIdepth,
  345. notostart - 2*ROIdepth,
  346. color='#d8b365',
  347. alpha=0.5
  348. )
  349. # Plot average expression profiles for each channel along AP axis for midline cells
  350. for channel in ['SOX2_01_norm', 'TBXT_01_norm',]:
  351. sns.lineplot(data=epi_nuc_ap_midline_avg, x='AP_bin', y=channel,
  352. ax=ax[0], color=colors.get(channel.split('_')[0], None),
  353. errorbar='sd')
  354. # turn off axis labels ax[0].set_xlabel('')
  355. ax[0].set_ylabel('')
  356. ax[0].set_title('')
  357. ax[0].set_yticks([0.2,0.4,0.6])
  358. # Plot ratio of TBXT to SOX2 along AP axis for midline cells
  359. sns.lineplot(data=epi_nuc_ap_midline_avg, x='AP_bin', y='TBXT_SOX2_ratio_log',
  360. ax=ax[1], color=colors.get('RATIO', None))
  361. ax[1].set_xlabel('')
  362. ax[1].set_ylabel('')
  363. # Looop over datasets (single genes, double genes, and ratio) to plot positional error along AP axis for midline cells
  364. for name, res in results_AP_midline.items():
  365. lower = res['lower']
  366. upper = res['upper']
  367. sigma = res['sigma']
  368. if name in ["SOX2", "TBXT"]:
  369. axs = ax[2]
  370. else:
  371. axs = ax[3]
  372. axs.plot(
  373. res["AP_bin"],
  374. sigma,
  375. label=name,
  376. color=colors.get(name, None)
  377. )
  378. # Calculate asymmetric error bars for confidence intervals - ensure that error bars do not extend below 0
  379. yerr_lower = np.maximum(
  380. 0,
  381. sigma - lower
  382. )
  383. yerr_upper = np.maximum(
  384. 0,
  385. upper - sigma
  386. )
  387. axs.errorbar(
  388. res["AP_bin"],
  389. sigma,
  390. yerr=[yerr_lower, yerr_upper],
  391. fmt='none',
  392. ecolor=colors.get(name, None),
  393. alpha=0.8,
  394. capsize=2,
  395. linewidth=1
  396. )
  397. # Set y-axis limites to the width of a ROI (10 cell widths or 50 microns)
  398. axs.set_ylim(0, 5)
  399. axs.set_xlim(-100, 50)
  400. axs.set_yticks(np.arange(0, 6, 1), minor=True) # add minor ticks every 1 unit
  401. # put y = 5 line as single cell positional error
  402. axs.axhline(1, color='black', linestyle='--', alpha=0.5, linewidth=1)
  403. # set all axis labels as size 8
  404. for axs in ax:
  405. axs.set_xlabel('')
  406. axs.set_ylabel('')
  407. axs.tick_params(axis='x', labelsize=8)
  408. axs.tick_params(axis='y', labelsize=8)
  409. plt.tight_layout( h_pad=0.1, w_pad=0.1)
  410. plt.savefig(os.path.join('Figures', 'AP_midline_positional_information.pdf'))
  411. plt.show()
  412. # =========== Now perform same analysis for medial-lateral axis ===========
  413. results_medial_lateral = {}
  414. for name, cols in genesets.items():
  415. data, positions, embryos = make_data_array(
  416. epi_nuc_medial_lateral_avg,
  417. cols,
  418. bin_col="LR_bin"
  419. )
  420. sigma= (
  421. positional_error(
  422. data,
  423. smoothing=smoothing_internal
  424. )
  425. )
  426. # calculate SEM as std/sqrt(n) where n is the number of shuffles (nshuffles)
  427. nshuffles = sigma.shape[0]
  428. mean_sigma = np.nanmean(sigma, axis=0)
  429. sem_sigma = np.nanstd(sigma, axis=0) / np.sqrt(nshuffles)
  430. lower = mean_sigma - 1.96 * sem_sigma
  431. upper = mean_sigma + 1.96 * sem_sigma
  432. results_medial_lateral[name] = {
  433. "LR_bin": positions,
  434. "sigma": mean_sigma,
  435. "lower": lower,
  436. "upper": upper
  437. }
  438. ROIdepth = 50
  439. fig, ax = plt.subplots(2,2, figsize=(3, 2), sharex=True)
  440. ax = ax.flatten()
  441. colors = {
  442. "SOX2": "#D65BC6",
  443. "TBXT": "#28b834",
  444. "SOX2_TBXT": "#000000",
  445. "RATIO": "#db4444"
  446. }
  447. for axs in ax:
  448. axs.axvspan(
  449. -ROIdepth/2,
  450. ROIdepth/2,
  451. color='#5ab4ac',
  452. alpha=0.5
  453. )
  454. axs.axvspan(
  455. -ROIdepth*1.5,
  456. -ROIdepth/2,
  457. color="#808080",
  458. alpha=0.5
  459. )
  460. axs.axvspan(
  461. ROIdepth*1.5,
  462. ROIdepth/2,
  463. color="#808080",
  464. alpha=0.5
  465. )
  466. axs.axvspan(
  467. -ROIdepth*1.5,
  468. -ROIdepth*2.5,
  469. color='#d8b365',
  470. alpha=0.5
  471. )
  472. axs.axvspan(
  473. ROIdepth*1.5,
  474. ROIdepth*2.5,
  475. color='#d8b365',
  476. alpha=0.5
  477. )
  478. for channel in ['SOX2_01_norm', 'TBXT_01_norm']:
  479. sns.lineplot(data=epi_nuc_medial_lateral_avg, x='LR_bin', y=channel,
  480. ax=ax[0], color=colors.get(channel.split('_')[0], None),
  481. errorbar='sd')
  482. ax[0].set_ylabel('')
  483. ax[0].set_title('')
  484. sns.lineplot(data=epi_nuc_medial_lateral_avg, x='LR_bin', y='TBXT_SOX2_ratio_log',
  485. ax=ax[1], color=colors.get('RATIO', None))
  486. ax[1].set_xlabel('')
  487. ax[1].set_ylabel('')
  488. ax[1].set_yticks([-0.2,0, 0.2])
  489. ax[1].set_ylim(-0.35, 0.35)
  490. for name, res in results_medial_lateral.items():
  491. if name in ["SOX2", "TBXT"]:
  492. axs = ax[2]
  493. else:
  494. axs = ax[3]
  495. axs.plot(
  496. res["LR_bin"],
  497. res["sigma"],
  498. label=name,
  499. color=colors.get(name, None)
  500. )
  501. yerr_lower = np.maximum(
  502. 0,
  503. res["sigma"] - res["lower"]
  504. )
  505. yerr_upper = np.maximum(
  506. 0,
  507. res["upper"] - res["sigma"]
  508. )
  509. axs.errorbar(
  510. res["LR_bin"],
  511. res["sigma"],
  512. yerr=[yerr_lower, yerr_upper],
  513. fmt='none',
  514. ecolor=colors.get(name, None),
  515. alpha=0.8,
  516. capsize=2,
  517. linewidth=1
  518. )
  519. axs.set_ylim(0, 5)
  520. axs.set_xlim(-125, 125)
  521. axs.set_yticks(np.arange(0, 6, 1), minor=True)
  522. # put y = 5 line as single cell positional error
  523. axs.axhline(1, color='black', linestyle='--', alpha=0.5, linewidth=1)
  524. # set all axis labels as size 8
  525. for axs in ax:
  526. axs.set_xlabel('')
  527. axs.set_ylabel('')
  528. axs.tick_params(axis='x', labelsize=8)
  529. axs.tick_params(axis='y', labelsize=8)
  530. plt.tight_layout( h_pad=0.1, w_pad=0.1)
  531. plt.savefig(os.path.join('Figures', 'St1_ML_positional_information.pdf'))
  532. plt.show()

Figure 6E.ipynb at commit cfcfa4e, under MIT · at the source

Overview

Authors: Anahí Binagui-Casas1, Anna Granés1, Alberto S. Ceccarelli2, Matthew French3, Filip J. Wymeersch4, Rosa Portero Migueles1, Jennifer Annoh1, Yali Huang5, Eleni P. Karagianni6, Frederick C. K. Wong7, Raffee Wright1, Anna Sophie Brumm1, Daniel Lopez Ramajo8, Minoru Takasato4,9, Sally Lowell1, Osvaldo Chara2,10, Valerie Wilson1
  1. Institute for Regeneration and Repair, Centre for Regenerative Medicine, Institute for Stem Cell Research, School of Biological Sciences, University of Edinburgh, Edinburgh, United Kingdom
  2. School of Biosciences, University of Nottingham, Sutton Bonington Campus, Nottingham, United Kingdom
  3. Department of Genetics, University of Cambridge, Cambridge, United Kingdom
  4. RIKEN Center for Biosystems Dynamics Research, Kobe, Japan
  5. Department of Cell, Developmental and Integrative Biology, University of Alabama at Birmingham, Birmingham, Alabama, United States of America
  6. GSK, Stevenage, United Kingdom
  7. The Wellcome Sanger Institute, Wellcome Genome Campus, Cambridge, United Kingdom
  8. Immunology Unit, Department of Pathology and Experimental Therapy, School of Medicine, Universitat de Barcelona, Barcelona, Spain
  9. Laboratory of Molecular Cell Biology and Development, Department of Animal Development and Physiology, Graduate School of Biostudies, Kyoto University, Kyoto, Japan
  10. Instituto de Tecnología, Universidad Argentina de la Empresa, Buenos Aires, Argentina
Journal: PLoS biology, volume 24, issue 8, article e3003960
Dates: received 5 December 2025; accepted 6 August 2026; published online 25 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1371/journal.pbio.3003960 · PMID 42640957 · PMCID PMC13533433 · OpenAlex W4410449633
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: mouse (organism)
Methods: Statistics, Evoked potentials, Connectivity
MeSH: Fetal Proteins*, Mesoderm*, Neural Stem Cells*, T-Box Domain Proteins*, Animals, Brachyury Protein, Cell Differentiation, Cell Line, Cell Lineage, Gene Expression Regulation, Developmental, Mice, Mouse Embryonic Stem Cells, SOXB1 Transcription Factors (* major topic)
Journal subjects: Biology and Life Sciences, Developmental Biology, Cell Differentiation, Embryology, Mesoderm, Genetics, Gene Expression, Embryos, Molecular Biology, Molecular Biology Techniques, Cloning, Research and Analysis Methods, Spectrum Analysis Techniques, Spectrophotometry, Cytophotometry, Flow Cytometry, Cell Biology, Signal Transduction, Cell Signaling, Notch Signaling, Organism Development, Organogenesis, Somites
Topic: Cellular Mechanics and Interactions (Cell Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: RCUK | Medical Research Council (MRC) (MR/K011200, MR/S008799/1); Carnegie Trust for the Universities of Scotland (RIG012666); Agencia Nacional de Promoción Científica y Tecnológica (ANPCyT) (PICT-2019-03828); RCUK | Biotechnology and Biological Sciences Research Council (BBSRC) (BB/X014908/1, BB/T00875X/1); UADE (A23T01, P26T02); Wellcome Trust (220298); Japan Society for the Promotion of Science (JP19K16157); Erasmus+; British society for developmental biology (Gurdon summer studentship)
Citations: not cited yet (Europe PMC); 63 references in the paper

Abstract

Neuromesodermal progenitors (NMPs) produce the spinal cord and musculoskeleton in the elongating anterior-posterior axis. In vivo, NMPs possess dual potency, coinciding with regions co-expressing SOX2 and Brachyury (TBXT). In vitro, SOX2/TBXT co-expressing cells can be produced from pluripotent cells and, like their in vivo counterparts, can produce neural tube and somitic mesoderm. However, the functional characteristics of in vitro SOX2/TBXT co-expressing cells remain unclear, confounding comparisons with in vivo data. To address this, we developed a dual Sox2/Tbxt reporter mouse ESC line. SOX2/TBXT reporter-positive cells emerge in vitro from pluripotent populations with dynamics that mirror their appearance in the embryo. Purified SOX2/TBXT co-expressing populations can differentiate towards neurectoderm or mesoderm, including lateral mesoderm upon BMP stimulation. In gastruloids, quantitative live imaging shows that WNT or NOTCH inhibition rapidly leads to downregulation of TBXT expression and diminished axial extension. We show that clonally plated SOX2/TBXT co-expressing cells are bipotent NMPs that can also self-propagate. By combining clonal analysis with mathematical inference, we identify two thresholds of TBXT and/or SOX2 expression, switching clonal output from neural- to mesoderm-biased, and from mesoderm-biased to mesoderm-specified. Image analysis of embryonic NMPs supports a model whereby SOX2 and TBXT independently influence neuromesodermal differentiation. Thus, this Sox2/Tbxt double reporter cell line highlights unsuspected heterogeneity in NMPs, and together with image analysis of embryonic SOX2/TBXT levels, challenges the assumption that neuromesodermal fate choice is primarily governed by mutual antagonism between SOX2/TBXT.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.

Zenodo 21650680

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the supplementary material
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file), SciPy (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
4 files

Zenodo 21676159

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data Availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)

Zenodo 15802710

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data Availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)

Zenodo 17192095

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the text, “Gastruloid culture, live imaging and quantitativ”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)

Zenodo 17154933

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the text, “Mathematical inference for in vitro data”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file), SciPy (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
4 files

mattfrenchh/pringle

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 20874e6bf622182dba48f4d516c490805c577fa0, 29 June 2025
Languages: Jupyter (8), R (5), MATLAB (4), Python (1)
Size: 38 files, 18 scripts
Software Heritage: not archived
Found in: the Zenodo archive record
Holds: README, license file, 8 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (9 files), NumPy (9 files), pandas (9 files), SciPy (9 files), scikit-image (6 files), ggplot2 (5 files), patchwork (5 files), reshape2 (3 files), scikit-learn (3 files), Statistics and Machine Learning Toolbox (2 files), tidyverse (2 files), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
20 files

albertoceccarelli/brachyury-expression-levels-predict-lineage-potential

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: bc9773d9a9c5ccd461e3e80ae7152c3c45043898, 28 July 2026
Languages: Jupyter (2)
Size: 5 files, 2 scripts
Software Heritage: not archived
Found in: the Zenodo archive record
Holds: README, license file, 2 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file), SciPy (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
4 files

mattfrenchh/binagui-casas_granes_et.al._2026_fig4_6_s5_s6

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: cfcfa4ec945aac281ec23c367d90b362afb4e95f, 29 July 2026
Languages: Jupyter (3), Shell (3), Python (3)
Size: 15 files, 9 scripts
Software Heritage: not archived
Found in: the Zenodo archive record
Holds: README, license file, 3 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (6 files), NumPy (6 files), pandas (5 files), scikit-image (4 files), tifffile (4 files), SciPy (3 files), seaborn (3 files), OpenCV (1 file), scikit-learn (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
11 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 8 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 33 scripts, each with its path and the digest of its content;
  • 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability

All relevant data can be found within the article, supplementary information and repositories as stated below. RNAseq raw data can be found at GEO accession number GSE339040. FCS files and data related to flow cytometry experiments and gene expression analyses can be found in Zenodo DOI: https://doi.org/10.5281/zenodo.21708690. Code and data for Figs 5E, 5F, S3G–S3I and S4 can be found in Zenodo DOI: https://doi.org/10.5281/zenodo.21650680. Code and data for Figs 4, 6, S5 and S6 can be found in Zenodo DOIs: https://doi.org/10.5281/zenodo.21676159 and https://doi.org/10.5281/zenodo.15802710.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 17 authors, 13 MeSH terms, 9 funders, 61 references.

Cite

This paper

Binagui-Casas, A., Granés, A., Ceccarelli, A. S., French, M., Wymeersch, F. J., Portero Migueles, R., Annoh, J., Huang, Y., Karagianni, E. P., Wong, F. C. K., Wright, R., Brumm, A. S., Lopez Ramajo, D., Takasato, M., Lowell, S., Chara, O., & Wilson, V. (2026). Brachyury expression levels predict lineage potential and axis-forming ability of in vitro-derived neuromesodermal progenitors. PLoS biology, 24(8), e3003960. https://doi.org/10.1371/journal.pbio.3003960

BibTeX

@article{binaguicasas2026brachyury,
author = {Binagui-Casas, Anahí and Granés, Anna and Ceccarelli, Alberto S. and French, Matthew and Wymeersch, Filip J. and Portero Migueles, Rosa and Annoh, Jennifer and Huang, Yali and Karagianni, Eleni P. and Wong, Frederick C. K. and Wright, Raffee and Brumm, Anna Sophie and Lopez Ramajo, Daniel and Takasato, Minoru and Lowell, Sally and Chara, Osvaldo and Wilson, Valerie},
title = {{Brachyury expression levels predict lineage potential and axis-forming ability of in vitro-derived neuromesodermal progenitors}},
journal = {PLoS biology},
year = {2026},
month = aug,
volume = {24},
number = {8},
pages = {e3003960},
publisher = {PLOS},
issn = {1544-9173},
doi = {10.1371/journal.pbio.3003960},
url = {https://doi.org/10.1371/journal.pbio.3003960},
pmid = {42640957},
pmcid = {PMC13533433}
}

RIS

TY - JOUR
AU - Binagui-Casas, Anahí
AU - Granés, Anna
AU - Ceccarelli, Alberto S.
AU - French, Matthew
AU - Wymeersch, Filip J.
AU - Portero Migueles, Rosa
AU - Annoh, Jennifer
AU - Huang, Yali
AU - Karagianni, Eleni P.
AU - Wong, Frederick C. K.
AU - Wright, Raffee
AU - Brumm, Anna Sophie
AU - Lopez Ramajo, Daniel
AU - Takasato, Minoru
AU - Lowell, Sally
AU - Chara, Osvaldo
AU - Wilson, Valerie
TI - Brachyury expression levels predict lineage potential and axis-forming ability of in vitro-derived neuromesodermal progenitors
T2 - PLoS biology
J2 - PLoS Biol
PY - 2026
DA - 2026/08/25
VL - 24
IS - 8
SP - e3003960
SN - 1544-9173
PB - PLOS
DO - 10.1371/journal.pbio.3003960
UR - https://doi.org/10.1371/journal.pbio.3003960
LA - en
ER -

CSL-JSON

{
"id": "10.1371/journal.pbio.3003960",
"type": "article-journal",
"title": "Brachyury expression levels predict lineage potential and axis-forming ability of in vitro-derived neuromesodermal progenitors",
"container-title": "PLoS biology",
"author": [
{
"family": "Binagui-Casas",
"given": "Anahí"
},
{
"family": "Granés",
"given": "Anna"
},
{
"family": "Ceccarelli",
"given": "Alberto S."
},
{
"family": "French",
"given": "Matthew"
},
{
"family": "Wymeersch",
"given": "Filip J."
},
{
"family": "Portero Migueles",
"given": "Rosa"
},
{
"family": "Annoh",
"given": "Jennifer"
},
{
"family": "Huang",
"given": "Yali"
},
{
"family": "Karagianni",
"given": "Eleni P."
},
{
"family": "Wong",
"given": "Frederick C. K."
},
{
"family": "Wright",
"given": "Raffee"
},
{
"family": "Brumm",
"given": "Anna Sophie"
},
{
"family": "Lopez Ramajo",
"given": "Daniel"
},
{
"family": "Takasato",
"given": "Minoru"
},
{
"family": "Lowell",
"given": "Sally"
},
{
"family": "Chara",
"given": "Osvaldo"
},
{
"family": "Wilson",
"given": "Valerie"
}
],
"container-title-short": "PLoS Biol",
"volume": "24",
"issue": "8",
"page": "e3003960",
"DOI": "10.1371/journal.pbio.3003960",
"PMID": "42640957",
"PMCID": "PMC13533433",
"ISSN": "1544-9173",
"publisher": "PLOS",
"URL": "https://doi.org/10.1371/journal.pbio.3003960",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
25
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: OpenCV, scikit-image, reshape2, 11 other tools, mouse
[2] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: tifffile, OpenCV, scikit-image, 10 other tools
[3] doi:10.1186/s11689-026-09713-0 [code]
DRP1 mutations associated with EMPF1 encephalopathy perturb the transcriptional profile and maturation of cortical neurons.
Journal: Journal of neurodevelopmental disorders
In common: tifffile, OpenCV, scikit-image, 9 other tools
[4] doi:10.1371/journal.pcbi.1014571 [code]
SynAPSeg: A novel dataset and image analysis framework for deep learning-based synapse detection and quantification.
Journal: PLoS computational biology
In common: tifffile, OpenCV, scikit-image, 7 other tools, 2 references
[5] doi:10.1002/imt2.70163 [code]
Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.
Journal: iMeta
In common: OpenCV, scikit-image, reshape2, 9 other tools, mouse
[6] doi:10.1093/nc/niag029 [code]
A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.
Journal: Neuroscience of consciousness
In common: scikit-image, patchwork, statsmodels, 9 other tools, 1 reference
[7] doi:10.1038/s41467-026-71458-0 [code]
Early differential impact of MeCP2 mutations on functional networks in Rett syndrome patient-derived human cortical organoids.
Journal: Nature communications
In common: tifffile, OpenCV, scikit-image, 8 other tools
[8] doi:10.7554/elife.109717 [code]
Retrosplenial cortex enables context-dependent goal-directed sensorimotor transformation.
Journal: eLife
In common: tifffile, OpenCV, scikit-image, 8 other tools, mouse
[9] doi:10.1002/advs.202522762 [code]
Enhancing Maturation of Human Neuromuscular Organoids via Electrical Stimulation.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: tifffile, scikit-image, patchwork, 7 other tools, 1 reference
[10] doi:10.1073/pnas.2609132123 [code]
A human lysosomal storage disorder toolkit for decoding proteome landscapes in cortical-like and dopaminergic-like induced neurons.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: tifffile, scikit-image, reshape2, 8 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.