OSCR

Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Experimental Procedures › Bioinformatic and Computational Analysis ↔ python/PCA and Volcano plot.ipynb, lines 1063–1097 · score 0.79 · gProfiler, CC, KEGG, MF, REAC, WP
  2. [2] § Experimental Procedures › Bioinformatic and Computational Analysis ↔ python/2025-ipscs-scp-github.ipynb, lines 220–223 · score 0.61 · UMAP, Scanpy, pp, resolution, neighbors, Leiden
  3. [3] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/Cluster heatmap and centroid distances.ipynb, lines 298–330 · score 0.60 · highlights small distances, pairwise Euclidean distances, clustering heatmap, proteomic
  4. [4] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/PCA and Volcano plot.ipynb, lines 718–744 · score 0.59 · ALDH1L1, CD44, MKI67, TUBB3, SOX2, SYN1
  5. [5] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/PCA and Volcano plot.ipynb, lines 1177–1261 · score 0.56 · GO term, gliogenesis, glutamatergic, astrocytes, volcano, enrichment
  6. [6] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/PCA and Volcano plot.ipynb, lines 1389–1457 · score 0.55 · Log2 fold change, Mann Whitney, NOTCH1, Volcano, bar, astrocytes
  7. [7] § Experimental Procedures › Linear Modeling ↔ python/Baseline linear model .ipynb, lines 120–156 · score 0.52 · linear model, OLS, ANOVA, fitted, row, Log2

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 1,620 lines · 49 KB · no license · 4 matches

  1. # %%
  2. import pandas as pd
  3. import plotly.express as px
  4. from sklearn.decomposition import PCA
  5. from sklearn.impute import KNNImputer
  6. import numpy as np
  7. import matplotlib.pyplot as plt
  8. import pandas as pd
  9. import seaborn as sns
  10. import matplotlib.pyplot as plt
  11. from scipy.spatial.distance import pdist, squareform
  12. # %%
  13. import os
  14. os.getcwd()
  15. # %%
  16. # Load your spreadsheet data into a Pandas DataFrame
  17. df = pd.read_csv(r"..\\data\2025\dox-mzml.pg_matrix (1).tsv", sep='\t', index_col=2)
  18. df
  19. df.drop(['Protein.Group', 'Protein.Names', 'First.Protein.Description','N.Sequences','N.Proteotypic.Sequences'],axis=1, inplace=True)
  20. df
  21. # %%
  22. df.columns = [col.split("\\")[-1].split(".")[0] for col in df.columns]
  23. df
  24. # %%
  25. np.log(df).boxplot(fontsize=5)
  26. #plt.axhline(df.count(axis=0), color='red', linestyle='--', label='Mean')
  27. plt.title('raw log2 quantities per sample')
  28. #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment/boxquant_raw.svg", bbox_inches='tight')
  29. plt.show()
  30. # %%
  31. df.count(axis=0)
  32. # %%
  33. plt.rcParams['figure.figsize'] = 3,3
  34. #plt.rcParams['axes.grid'] = False
  35. df.count(axis=0).hist(bins=20, grid=False)
  36. plt.axvline(x=7500, color='black', linewidth=3)
  37. plt.xlabel('number of protein IDs')
  38. plt.ylabel('count')
  39. #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment/proteins_per_condition.svg", bbox_inches='tight')
  40. plt.show()
  41. # %%
  42. df.T.index
  43. # %%
  44. dropthese = df.T.index[df.count()<7500]
  45. dropthese
  46. # %%
  47. df = df.drop(dropthese, axis=1)
  48. # %%
  49. df.count()
  50. # %%
  51. # Assuming your DataFrame is named 'df'
  52. column_sums = df.sum(axis=0) # Calculate the sum of each column
  53. # Divide each element by the corresponding column sum
  54. df_normalized = df.div(column_sums, axis=1)
  55. df_normalized
  56. # %%
  57. # Handle missing values (e.g., impute with mean or median), imputation may contain very small numbers
  58. # Create a KNN imputer object
  59. imputer = KNNImputer(n_neighbors=5) # Adjust n_neighbors as needed
  60. # Impute missing values using KNN
  61. df_imputed = imputer.fit_transform(df_normalized)
  62. # Create a new DataFrame with imputed values
  63. df_imputed = pd.DataFrame(df_imputed, columns=df.columns, index= df.index)
  64. df_imputed*10e6
  65. # %%
  66. #[0,:] all columns in the first row, what type of distribution
  67. df_imputed.iloc[0,:].hist()
  68. # Set the x and y labels
  69. plt.xlabel('Value') # Label for x-axis
  70. plt.ylabel('Frequency') # Label for y-axis
  71. plt.show()
  72. # %%
  73. df_imputed = np.log(df_imputed*10e6) #np.log (...) applies the natural algorithm (log base e) to all values, improve linearity for PCA
  74. #make data more normally distributed and reduce the effects of outliers
  75. # %%
  76. df_imputed.iloc[0,:].hist()
  77. plt.xlabel('Value') # Label for x-axis
  78. plt.ylabel('Frequency') # Label for y-axis
  79. plt.show()
  80. # %%
  81. # Extract group labels from column names
  82. # Jessi code=group_labels = [int(col[-1]) for col in df.columns]
  83. #My group_labels are in the firts row
  84. group_labels = [col[0] for col in df.columns]
  85. group_labels
  86. # %%
  87. #shape of the transposed values
  88. df_imputed.values.T.shape
  89. # %%
  90. # Transpose the DataFrame
  91. df_transposed = df_imputed.T #T stands for transpose it flips the data frame rows into columns and columns into rows
  92. #samples as rows are requiered for PCA and StandardScaler
  93. # %%
  94. # Print the new shape
  95. print(df_transposed.shape)
  96. # %%
  97. df_transposed
  98. # %%
  99. # Standardize features (per column)
  100. from sklearn.preprocessing import StandardScaler
  101. scaler = StandardScaler()
  102. df_scaled = scaler.fit_transform(df_imputed.values) #df_scaled standarized
  103. # %%
  104. # Reconstruct a DataFrame from df_scaled for easier plotting
  105. df_scaled_df = pd.DataFrame(df_scaled, index=df_imputed.index, columns=df_imputed.columns)
  106. # Boxplot per sample (i.e., each box is one sample)
  107. plt.figure(figsize=(14, 6))
  108. df_scaled_df.boxplot(showfliers=False) # Set to True if you want to show outliers
  109. # Plot formatting
  110. plt.title("Boxplot of Standardized Feature Values per Sample", fontsize=16)
  111. plt.xlabel("Samples", fontsize=14)
  112. plt.ylabel("Standardized Feature Values", fontsize=14)
  113. plt.xticks(rotation=90, fontsize=10)
  114. plt.yticks(fontsize=12)
  115. plt.tight_layout()
  116. plt.show()
  117. # %%
  118. print("Shape of df_scaled_df:", df_scaled_df.shape)
  119. # %%
  120. print("Any missing values?", df_scaled_df.isnull().values.any())
  121. # %%
  122. print("Overall mean:", df_scaled_df.values.mean())
  123. print("Overall std:", df_scaled_df.values.std())
  124. # %%
  125. df_scaled_T = df_scaled_df.T # Now shape is (37, 9782)
  126. # %%
  127. print("Shape of df_scaled_df:", df_scaled_df.T.shape)
  128. # %%
  129. print("Any missing values?", df_scaled_df.T.isnull().values.any())
  130. # %%
  131. print("Overall mean:", df_scaled_df.T.values.mean())
  132. print("Overall std:", df_scaled_df.T.values.std())
  133. # %%
  134. # Transpose so samples are rows
  135. df_scaled_T = df_scaled_df.T
  136. # ✅ Convert all column names (features) to strings
  137. df_scaled_T.columns = df_scaled_T.columns.astype(str)
  138. # PCA
  139. from sklearn.decomposition import PCA
  140. pca = PCA(n_components=2)
  141. pca_data = pca.fit_transform(df_scaled_T)
  142. # Step 2: Create DataFrame with PCA results
  143. pca_df = pd.DataFrame(pca_data, columns=["PC1", "PC2"])
  144. pca_df["Sample"] = df_scaled_T.index # Sample names from rows
  145. pca_df["Group"] = pca_df["Sample"].str.split("_").str[0] # Customize based on your naming # not good for
  146. # Step 3: Plot with Matplotlib
  147. fig, ax = plt.subplots(figsize=(10, 8))
  148. # Color by group
  149. for group in pca_df["Group"].unique():
  150. group_data = pca_df[pca_df["Group"] == group]
  151. ax.scatter(group_data["PC1"], group_data["PC2"], label=f"{group}", s=150, alpha=0.7)
  152. # Optional: annotate each sample
  153. for _, row in pca_df.iterrows():
  154. ax.text(row["PC1"] + 0.3, row["PC2"] + 0.3, row["Sample"], fontsize=9, alpha=0.8)
  155. # Variance explained for axis labels
  156. explained_var = PCA(n_components=2).fit(df_scaled_T).explained_variance_ratio_ * 100
  157. ax.set_xlabel(f"PC1 ({explained_var[0]:.1f}%)", fontsize=14)
  158. ax.set_ylabel(f"PC2 ({explained_var[1]:.1f}%)", fontsize=14)
  159. # Final formatting
  160. ax.set_title("PCA Plot (Matplotlib)", fontsize=16, fontweight="bold")
  161. ax.legend(title="Group", fontsize=12)
  162. ax.tick_params(labelsize=12)
  163. plt.tight_layout()
  164. plt.show()
  165. # %%
  166. # Transpose so samples are rows
  167. df_scaled_T = df_scaled_df.T
  168. # ✅ Convert all column names (features) to strings
  169. df_scaled_T.columns = df_scaled_T.columns.astype(str)
  170. # PCA
  171. from sklearn.decomposition import PCA
  172. pca = PCA(n_components=2)
  173. pca_data = pca.fit_transform(df_scaled_T)
  174. # Step 2: Create DataFrame with PCA results
  175. pca_df = pd.DataFrame(pca_data, columns=["PC1", "PC2"])
  176. pca_df["Sample"] = df_scaled_T.index # Sample names from rows
  177. pca_df["Group"] = pca_df["Sample"].str[0] # Customize based on your naming # not good for
  178. # Step 3: Plot with Matplotlib
  179. fig, ax = plt.subplots(figsize=(12, 10))
  180. # Color by group
  181. for group in pca_df["Group"].unique():
  182. group_data = pca_df[pca_df["Group"] == group]
  183. ax.scatter(group_data["PC1"], group_data["PC2"], label=f"{group}", s=300, alpha=0.7) #dot size s=400
  184. # Optional: annotate each sample
  185. for _, row in pca_df.iterrows():
  186. ax.text(row["PC1"] + 0.3, row["PC2"] + 0.3, row["Sample"], fontsize=9, alpha=0.8)
  187. # Variance explained for axis labels
  188. explained_var = PCA(n_components=2).fit(df_scaled_T).explained_variance_ratio_ * 100
  189. ax.set_xlabel(f"PC1 ({explained_var[0]:.1f}%)", fontsize=14)
  190. ax.set_ylabel(f"PC2 ({explained_var[1]:.1f}%)", fontsize=14)
  191. # Final formatting
  192. ax.set_title("PCA Plot (Matplotlib)", fontsize=16, fontweight="bold") #axis label fontsize+
  193. ax.legend(title="Group", fontsize=12)
  194. ax.tick_params(labelsize=14)
  195. plt.tight_layout()
  196. plt.show()
  197. # %%
  198. # Step 1: Perform PCA
  199. pca = PCA(n_components=2)
  200. pca_data = pca.fit_transform(df_scaled_T)
  201. # Step 2: Create DataFrame with PCA results
  202. pca_df = pd.DataFrame(pca_data, columns=["PC1", "PC2"])
  203. pca_df["Sample"] = df_scaled_T.index
  204. pca_df["Group"] = pca_df["Sample"].str[0] # Adjust if needed
  205. # Step 3: Plot
  206. fig, ax = plt.subplots(figsize=(12, 10))
  207. # Plot each group with larger dots
  208. for group in pca_df["Group"].unique():
  209. group_data = pca_df[pca_df["Group"] == group]
  210. ax.scatter(
  211. group_data["PC1"],
  212. group_data["PC2"],
  213. label=f"Group {group}",
  214. s=700,
  215. alpha=0.7
  216. )
  217. # Axis labels with % variance explained
  218. explained = pca.explained_variance_ratio_ * 100
  219. #ax.set_title("PCA Plot", fontsize=20, fontweight='bold')
  220. ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=16)
  221. ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=16)
  222. # Make axis lines (spines) thicker
  223. ax.spines['bottom'].set_linewidth(2)
  224. ax.spines['left'].set_linewidth(2)
  225. ax.spines['top'].set_linewidth(2)
  226. ax.spines['right'].set_linewidth(2)
  227. # Bigger tick labels
  228. ax.tick_params(width=2, length=6, labelsize=25)
  229. # Legend
  230. ax.legend(title="Group", fontsize=12, title_fontsize=14, bbox_to_anchor=(1.05, 1), loc='upper left')
  231. # Tight layout
  232. plt.tight_layout()
  233. # Step 4: Export as SVG
  234. #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\PCA_final.svg", format="svg")
  235. # Optional: Show plot
  236. plt.show()
  237. # %%
  238. from sklearn.decomposition import PCA
  239. pca = PCA(n_components=2)
  240. pca_data = pca.fit_transform(df_scaled_T)
  241. explained = pca.explained_variance_ratio_ * 100
  242. print("Explained variance:", pca.explained_variance_ratio_)
  243. # %%
  244. print("Groups detected:", pca_df["Group"].unique())
  245. print("Number of groups:", pca_df["Group"].nunique())
  246. # %%
  247. from sklearn.cluster import KMeans
  248. # Step 1: Cluster PCA data (2D) into 2 clusters
  249. kmeans = KMeans(n_clusters=2, random_state=42)
  250. pca_df["Cluster"] = kmeans.fit_predict(pca_df[["PC1", "PC2"]])
  251. # Step 2: Plot with cluster colors instead of groups
  252. fig, ax = plt.subplots(figsize=(10, 8))
  253. # Define cluster colors
  254. colors = ['red', 'blue'] # Or use more for more clusters
  255. for cluster in pca_df["Cluster"].unique():
  256. cluster_data = pca_df[pca_df["Cluster"] == cluster]
  257. ax.scatter(cluster_data["PC1"], cluster_data["PC2"],
  258. label=f"Cluster {cluster+1}", c=colors[cluster],
  259. s=200, alpha=0.7)
  260. # Axes and legend
  261. explained = pca.explained_variance_ratio_ * 100
  262. ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
  263. ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
  264. ax.set_title("PCA with KMeans Clustering", fontsize=16, fontweight="bold")
  265. ax.legend(title="KMeans Cluster", fontsize=12)
  266. ax.tick_params(labelsize=12)
  267. # Optional: Thicker axes
  268. for spine in ax.spines.values():
  269. spine.set_linewidth(2)
  270. ax.tick_params(width=2, length=6)
  271. plt.tight_layout()
  272. plt.show()
  273. # %%
  274. pd.crosstab(pca_df["Group"], pca_df["Cluster"])
  275. # %%
  276. from sklearn.cluster import KMeans
  277. import matplotlib.pyplot as plt
  278. # Step 1: Cluster PCA data (2 clusters for example)
  279. kmeans = KMeans(n_clusters=2, random_state=42)
  280. pca_df["Cluster"] = kmeans.fit_predict(pca_df[["PC1", "PC2"]])
  281. # Step 2: Get centroids from the KMeans model
  282. centroids = kmeans.cluster_centers_ # shape: (n_clusters, 2)
  283. # Step 3: Plot
  284. fig, ax = plt.subplots(figsize=(10, 8))
  285. # Cluster colors
  286. colors = ['lightblue', 'gray']
  287. for cluster in pca_df["Cluster"].unique():
  288. cluster_data = pca_df[pca_df["Cluster"] == cluster]
  289. ax.scatter(
  290. cluster_data["PC1"],
  291. cluster_data["PC2"],
  292. label=f"Cluster {cluster+1}",
  293. c=colors[cluster],
  294. s=500,
  295. alpha=0.7
  296. )
  297. # Step 4: Plot centroids
  298. ax.scatter(
  299. centroids[:, 0], # PC1
  300. centroids[:, 1], # PC2
  301. marker='X',
  302. s=300,
  303. c='orange',
  304. label='Centroid',
  305. edgecolor='white',
  306. linewidth=2
  307. )
  308. # Axis formatting
  309. explained = pca.explained_variance_ratio_ * 100
  310. ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=22)
  311. ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=22)
  312. ax.set_title("PCA with KMeans Clustering and Centroids", fontsize=18, fontweight="bold")
  313. ax.tick_params(labelsize=22)
  314. for spine in ax.spines.values():
  315. spine.set_linewidth(2)
  316. ax.tick_params(width=2, length=6)
  317. #ax.legend(title="Cluster", fontsize=16)
  318. plt.tight_layout()
  319. #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\PCA_withKMeans.svg", format="svg")
  320. plt.show()
  321. # %%
  322. import seaborn as sns
  323. import matplotlib.pyplot as plt
  324. # Step 1: Transpose to get samples as rows
  325. df_plot = df_imputed.T
  326. # Step 2: Add group info based on sample names
  327. df_plot["Group"] = df_plot.index.str[0]
  328. # Step 3: Choose gene of interest (make sure it's in the column names now)
  329. gene_of_interest = "PRPH" # Replace with a valid gene name from df_imputed.index
  330. # Step 4: Move the gene from index to column if necessary
  331. df_plot.columns = df_plot.columns.astype(str)
  332. if gene_of_interest not in df_plot.columns:
  333. df_plot[gene_of_interest] = df_imputed.loc[gene_of_interest].values
  334. # Step 5: Violin plot
  335. plt.figure(figsize=(10, 6))
  336. sns.violinplot(data=df_plot, x="Group", y=gene_of_interest, inner="box", palette="Set2")
  337. # Step 6: Plot styling
  338. plt.title(f"Expression of {gene_of_interest} Across Groups", fontsize=16)
  339. plt.xlabel("Group", fontsize=14)
  340. plt.ylabel("Expression Level", fontsize=14)
  341. plt.xticks(fontsize=12)
  342. plt.yticks(fontsize=12)
  343. plt.tight_layout()
  344. plt.show()
  345. # %%
  346. print(df_imputed.index[:10].tolist())
  347. # %%
  348. import seaborn as sns
  349. import matplotlib.pyplot as plt
  350. # Step 1: Transpose to get samples as rows
  351. df_plot = df_imputed.T
  352. # Step 2: Add group info based on sample names
  353. df_plot["Group"] = df_plot.index.str[0]
  354. # Step 3: Choose gene of interest (make sure it's in the column names now)
  355. gene_of_interest = "SYN1" # Replace with a valid gene name from df_imputed.index
  356. # Step 4: Move the gene from index to column if necessary
  357. df_plot.columns = df_plot.columns.astype(str)
  358. if gene_of_interest not in df_plot.columns:
  359. df_plot[gene_of_interest] = df_imputed.loc[gene_of_interest].values
  360. # Step 5: Violin plot
  361. plt.figure(figsize=(10, 6))
  362. sns.violinplot(data=df_plot, x="Group", y=gene_of_interest, inner="box", palette="Set2")
  363. # Step 6: Plot styling
  364. plt.title(f"Expression of {gene_of_interest} Across Groups", fontsize=16)
  365. plt.xlabel("Group", fontsize=14)
  366. plt.ylabel("Expression Level", fontsize=14)
  367. plt.xticks(fontsize=12)
  368. plt.yticks(fontsize=12)
  369. plt.tight_layout()
  370. plt.show()
  371. # %%
  372. import pandas as pd
  373. # Step 1: Transpose to get samples as rows
  374. df_plot = df_imputed.T.copy() # Now shape: (samples, genes)
  375. # Step 2: Add group info (based on sample name, e.g., 'A1' → group 'A')
  376. df_plot["Group"] = df_plot.index.str[0] # Adjust if your sample names are different
  377. # Step 3: Extract expression of the gene of interest (e.g., TP53)
  378. gene_of_interest = "SLC17A6" # Make sure it's a valid gene in df_imputed.index
  379. gene_data = df_plot[[gene_of_interest, "Group"]] # Keep only gene and group
  380. # Step 4: Export to Excel
  381. output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\SC17A6_expression_by_group.xlsx"
  382. #gene_data.to_excel(output_path)
  383. print("Exported:", output_path)
  384. # %% [markdown]
  385. # ### Differential expression between clusters.
  386. # %%
  387. # Step 1: Transpose df_imputed to get samples as rows
  388. df_expr = df_imputed.T.copy() # shape: (samples, proteins)
  389. # Step 2: Add Cluster labels to expression matrix
  390. df_expr["Cluster"] = pca_df.set_index("Sample").loc[df_expr.index, "Cluster"]
  391. # Step 3: Group by cluster and compute mean expression
  392. cluster_means = df_expr.groupby("Cluster").mean().T # shape: (proteins, 2 clusters)
  393. # Step 4: Compute absolute difference between clusters
  394. cluster_means["diff"] = abs(cluster_means[0] - cluster_means[1])
  395. # Step 5: Sort by difference and get top 10 proteins
  396. top_proteins = cluster_means.sort_values("diff", ascending=False).head(20)
  397. print("Top 20 proteins separating clusters:\n")
  398. print(top_proteins)
  399. # %%
  400. # Cross-tabulate clusters and original groups
  401. cluster_group_counts = pd.crosstab(pca_df["Cluster"], pca_df["Group"])
  402. # Display the mapping
  403. print(cluster_group_counts)
  404. # %%
  405. cluster_group_counts.T.plot(kind="bar", stacked=True, figsize=(10, 6), colormap="Set2")
  406. plt.title("Group Composition per Cluster")
  407. plt.xlabel("Group")
  408. plt.ylabel("Sample Count")
  409. plt.legend(title="KMeans Cluster")
  410. plt.tight_layout()
  411. plt.show()
  412. # %%
  413. import matplotlib.pyplot as plt
  414. import seaborn as sns
  415. import pandas as pd
  416. # Setup colors and marker shapes
  417. group_colors = {
  418. "A": "forestgreen", "B": "red", "C": "dodgerblue", "D": "darkorange",
  419. "E": "gold", "F": "tomato", "G": "purple", "H": "skyblue"
  420. }
  421. cluster_markers = {
  422. 0: "o", # circle
  423. 1: "s", # square
  424. }
  425. # Create figure
  426. fig, ax = plt.subplots(figsize=(12, 10))
  427. # Plot each sample by group (color) and cluster (shape)
  428. for cluster in sorted(pca_df["Cluster"].unique()):
  429. for group in sorted(pca_df["Group"].unique()):
  430. subset = pca_df[(pca_df["Cluster"] == cluster) & (pca_df["Group"] == group)]
  431. ax.scatter(
  432. subset["PC1"],
  433. subset["PC2"],
  434. color=group_colors.get(group, "gray"),
  435. marker=cluster_markers.get(cluster, "o"),
  436. s=200,
  437. alpha=0.8,
  438. label=f"Group {group}, Cluster {cluster}"
  439. )
  440. # Axis formatting
  441. explained = pca.explained_variance_ratio_ * 100
  442. ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
  443. ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
  444. ax.set_title("PCA: Color by Group, Shape by Cluster", fontsize=18, fontweight="bold")
  445. ax.tick_params(labelsize=12)
  446. # Legend (no duplicates)
  447. handles, labels = ax.get_legend_handles_labels()
  448. unique_labels = dict(zip(labels, handles))
  449. ax.legend(unique_labels.values(), unique_labels.keys(), title="Group + Cluster", bbox_to_anchor=(1.05, 1), loc='upper left')
  450. plt.tight_layout()
  451. plt.show()
  452. # %%
  453. import matplotlib.patches as mpatches
  454. import numpy as np
  455. fig, ax = plt.subplots(figsize=(12, 10))
  456. # Plot each group with unique color, marker shape based on cluster
  457. for cluster in sorted(pca_df["Cluster"].unique()):
  458. for group in sorted(pca_df["Group"].unique()):
  459. subset = pca_df[(pca_df["Group"] == group) & (pca_df["Cluster"] == cluster)]
  460. if subset.empty:
  461. continue
  462. # Plot points
  463. ax.scatter(
  464. subset["PC1"],
  465. subset["PC2"],
  466. color=group_colors.get(group, "gray"),
  467. marker=cluster_markers.get(cluster, "o"),
  468. s=200,
  469. alpha=0.7,
  470. label=f"{group}, Cluster {cluster}"
  471. )
  472. # Add ellipse around the group
  473. if len(subset) >= 3: # Need at least 3 points to define an ellipse
  474. cov = np.cov(subset[["PC1", "PC2"]].T)
  475. lambda_, v = np.linalg.eig(cov)
  476. lambda_ = np.sqrt(lambda_)
  477. ell = mpatches.Ellipse(
  478. xy=(subset["PC1"].mean(), subset["PC2"].mean()),
  479. width=lambda_[0] * 4, # 2 std dev
  480. height=lambda_[1] * 4,
  481. angle=np.rad2deg(np.arccos(v[0, 0])),
  482. edgecolor=group_colors.get(group, "gray"),
  483. facecolor='none',
  484. linewidth=2,
  485. linestyle="--",
  486. alpha=0.7
  487. )
  488. ax.add_patch(ell)
  489. # Axes
  490. explained = pca.explained_variance_ratio_ * 100
  491. ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
  492. ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
  493. ax.set_title("PCA with Ellipses by Group and Clustering Shape", fontsize=18)
  494. # Legend cleanup
  495. handles, labels = ax.get_legend_handles_labels()
  496. unique = dict(zip(labels, handles))
  497. ax.legend(unique.values(), unique.keys(), title="Group + Cluster", bbox_to_anchor=(1.05, 1), loc="upper left")
  498. plt.tight_layout()
  499. plt.show()
  500. # %%
  501. # Step 1: Remove Group A
  502. samples_to_keep = pca_df[pca_df["Group"] != "A"]["Sample"]
  503. df_filtered = df_imputed.loc[:, samples_to_keep]
  504. # Step 2: Transpose and fix column types
  505. df_filtered_T = df_filtered.T.copy()
  506. df_filtered_T.columns = df_filtered_T.columns.astype(str)
  507. # Step 3: Standardize
  508. from sklearn.preprocessing import StandardScaler
  509. scaler = StandardScaler()
  510. df_scaled_filtered = scaler.fit_transform(df_filtered_T)
  511. # ✅ Step 4: Run PCA on filtered data
  512. from sklearn.decomposition import PCA
  513. pca = PCA(n_components=2)
  514. pca_data_filtered = pca.fit_transform(df_scaled_filtered) # 🔥 This was missing
  515. # ✅ Step 5: Run KMeans clustering on PCA output
  516. from sklearn.cluster import KMeans
  517. kmeans = KMeans(n_clusters=2, random_state=42)
  518. clusters = kmeans.fit_predict(pca_data_filtered)
  519. # Step 6: Build new PCA DataFrame
  520. import pandas as pd
  521. pca_filtered_df = pd.DataFrame(pca_data_filtered, columns=["PC1", "PC2"])
  522. pca_filtered_df["Sample"] = df_filtered.columns
  523. pca_filtered_df["Group"] = pca_filtered_df["Sample"].str[0]
  524. pca_filtered_df["Cluster"] = clusters # ✅ Now this works!
  525. # Step 7: Plot
  526. import matplotlib.pyplot as plt
  527. fig, ax = plt.subplots(figsize=(10, 8))
  528. for cluster in pca_filtered_df["Cluster"].unique():
  529. cluster_data = pca_filtered_df[pca_filtered_df["Cluster"] == cluster]
  530. ax.scatter(
  531. cluster_data["PC1"], cluster_data["PC2"],
  532. label=f"Cluster {cluster}", s=200, alpha=0.7
  533. )
  534. explained = pca.explained_variance_ratio_ * 100
  535. ax.set_title("PCA after Removing Group A", fontsize=16)
  536. ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
  537. ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
  538. ax.legend(title="KMeans Cluster")
  539. plt.tight_layout()
  540. plt.show()
  541. # %%
  542. print(f"Explained variance (PC1): {explained[0]:.2f}%")
  543. print(f"Explained variance (PC2): {explained[1]:.2f}%")
  544. # %%
  545. from scipy.stats import zscore
  546. # Step 1: Extract gene expression across all samples
  547. gene = "NES"
  548. gene_expr = df_imputed.loc[gene] # This is a Series: index = sample names, values = expression
  549. # Step 2: Convert to DataFrame and add Cluster info
  550. gene_df = pd.DataFrame({"Expression": gene_expr})
  551. gene_df["Sample"] = gene_df.index
  552. gene_df["Cluster"] = gene_df["Sample"].map(pca_df.set_index("Sample")["Cluster"])
  553. # Step 3: Compute Z-score of expression across all samples
  554. gene_df["Z_score"] = zscore(gene_df["Expression"])
  555. # Step 4: Compare distributions across clusters
  556. import seaborn as sns
  557. import matplotlib.pyplot as plt
  558. plt.figure(figsize=(8, 6))
  559. sns.boxplot(data=gene_df, x="Cluster", y="Z_score", palette="Set2")
  560. sns.stripplot(data=gene_df, x="Cluster", y="Z_score", color="black", alpha=0.6)
  561. plt.title(f"Z-score of {gene} Expression Across Clusters", fontsize=16)
  562. plt.xlabel("KMeans Cluster", fontsize=14)
  563. plt.ylabel("Z-score", fontsize=14)
  564. plt.tight_layout()
  565. plt.show()
  566. # Step 4: Export to Excel
  567. output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\NES_zscores.xlsx"
  568. #gene_df.to_excel(output_path, index=False)
  569. print("✅ Exported Z-scores to:", output_path)
  570. # %%
  571. from scipy.stats import zscore
  572. # Step 1: Define your list of genes (or use top ones)
  573. genes_of_interest = ["NES", "TUBB3", "GFAP", "PRPH", "SLC17A7", "SLC17A6", "RBFOX3", "MAP2", "SYN1","SOX2","CD44", "SYP", "MKI67", "ALDH1L1", "DLG4", "DCX"] # Replace with your list
  574. # Step 2: Transpose df_imputed to shape (samples x genes)
  575. df_expr = df_imputed.T.copy()
  576. df_expr["Sample"] = df_expr.index
  577. df_expr["Group"] = df_expr["Sample"].str[0] # Adjust if needed
  578. df_expr["Cluster"] = df_expr["Sample"].map(pca_df.set_index("Sample")["Cluster"])
  579. # Step 3: Melt into long format
  580. df_long = df_expr[genes_of_interest + ["Sample", "Group", "Cluster"]].melt(
  581. id_vars=["Sample", "Group", "Cluster"],
  582. var_name="Gene",
  583. value_name="Expression"
  584. )
  585. # Step 4: Compute Z-scores for each gene across all samples
  586. df_long["Z_score"] = df_long.groupby("Gene")["Expression"].transform(zscore)
  587. # Step 5: Export to Excel
  588. output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\multi_gene_zscores1.xlsx"
  589. #df_long.to_excel(output_path, index=False)
  590. print("✅ Exported multi-gene Z-scores to:", output_path)
  591. # %%
  592. import seaborn as sns
  593. import matplotlib.pyplot as plt
  594. # Use the DataFrame from your earlier export, e.g., `df_long`
  595. plt.figure(figsize=(14, 8))
  596. sns.boxplot(data=df_long, x="Group", y="Z_score", hue="Gene", palette="Set2")
  597. plt.title("Z-score of Multiple Genes by Group", fontsize=16)
  598. plt.xlabel("Group", fontsize=14)
  599. plt.ylabel("Z-score", fontsize=14)
  600. plt.legend(title="Gene", bbox_to_anchor=(1.05, 1), loc="upper left")
  601. plt.tight_layout()
  602. plt.show()
  603. # %%
  604. plt.figure(figsize=(14, 8))
  605. sns.violinplot(
  606. data=df_long,
  607. x="Gene",
  608. y="Z_score",
  609. hue="Cluster", # ✅ Switch hue to cluster instead
  610. palette="Set2",
  611. dodge=True # Separates the clusters side-by-side
  612. )
  613. plt.title("Violin Plot of Gene Z-scores by Cluster", fontsize=16)
  614. plt.xlabel("Gene")
  615. plt.ylabel("Z-score")
  616. plt.legend(title="Cluster", bbox_to_anchor=(1.05, 1), loc="upper left")
  617. plt.tight_layout()
  618. plt.show()
  619. # %%
  620. from scipy.stats import mannwhitneyu
  621. import seaborn as sns
  622. import matplotlib.pyplot as plt
  623. import os
  624. # Create output folder
  625. #output_dir = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\violin_plots"
  626. #os.makedirs(output_dir, exist_ok=True)
  627. # ✅ Define custom colors (do this outside the loop — one time)
  628. custom_palette = {0: "lightblue", 1: "lightgray"}
  629. # Loop through genes
  630. for gene in df_long["Gene"].unique():
  631. plt.figure(figsize=(6, 5))
  632. # Subset data
  633. df_gene = df_long[df_long["Gene"] == gene]
  634. group0 = df_gene[df_gene["Cluster"] == 0]["Z_score"]
  635. group1 = df_gene[df_gene["Cluster"] == 1]["Z_score"]
  636. # Mann–Whitney U
  637. u_stat, p_val = mannwhitneyu(group0, group1, alternative='two-sided')
  638. # Significance stars
  639. if p_val < 0.001:
  640. sig = "***"
  641. elif p_val < 0.01:
  642. sig = "**"
  643. elif p_val < 0.05:
  644. sig = "*"
  645. else:
  646. sig = "ns"
  647. # ✅ Plot using your custom colors
  648. sns.violinplot(
  649. data=df_gene,
  650. x="Cluster",
  651. y="Z_score",
  652. hue="Cluster",
  653. split=True,
  654. palette=custom_palette,
  655. dodge=False
  656. )
  657. plt.legend().remove()
  658. # Style
  659. plt.title(f"{gene} (Mann–Whitney p = {p_val:.3e}, {sig})", fontsize=16)
  660. plt.xlabel("Cluster", fontsize=16)
  661. plt.ylabel("Z-score", fontsize=16)
  662. plt.xticks(fontsize=16)
  663. plt.yticks(fontsize=16)
  664. # Thicken axes lines
  665. ax = plt.gca()
  666. for spine in ax.spines.values():
  667. spine.set_linewidth(1)
  668. # Major ticks
  669. ax.tick_params(axis='both', which='major', labelsize=16, length=6, width=1.5)
  670. # Add significance annotation
  671. y_max = df_gene["Z_score"].max()
  672. plt.ylim(top=y_max + 1)
  673. plt.plot([0.25, 0.75], [y_max + 0.3]*2, lw=1.5, c='black')
  674. plt.text(0.5, y_max + 0.5, sig, ha='center', va='center', fontsize=16)
  675. # Save as SVG
  676. output_path = os.path.join(output_dir, f"{gene}_violin_cluster.svg")
  677. #plt.savefig(output_path, format="svg", bbox_inches="tight")
  678. plt.tight_layout()
  679. plt.show()
  680. # %% [markdown]
  681. # ## For your volcano plot, you want to compare expression values (not Z-scores) between clusters.
  682. #
  683. # ## So:
  684. # ## ❌ Don't use df_scaled_df or df_scaled_T for volcano
  685. # ## ✅ Use the log-transformed, imputed, normalized data → df_imputed
  686. # %%
  687. # Already log-transformed, normalized, and imputed
  688. df_expression = df_imputed
  689. # %%
  690. df_imputed
  691. # %%
  692. sample_to_group = df_expression.columns.str[0] # assumes sample names start with group letter
  693. # %%
  694. group1_labels = ['A', 'B', 'C', 'D']
  695. group2_labels = ['E', 'F', 'G', 'H']
  696. sample_groups = df_expression.columns.to_series().str[0]
  697. group1_samples = sample_groups[sample_groups.isin(group1_labels)].index
  698. group2_samples = sample_groups[sample_groups.isin(group2_labels)].index
  699. # %%
  700. group1 = df_expression[group1_samples]
  701. group2 = df_expression[group2_samples]
  702. # %%
  703. from scipy.stats import ttest_ind
  704. import numpy as np
  705. import pandas as pd
  706. from statsmodels.stats.multitest import multipletests
  707. import matplotlib.pyplot as plt
  708. # Calculate means
  709. mean_group1 = group1.mean(axis=1)
  710. mean_group2 = group2.mean(axis=1)
  711. # Log2 fold change (already log-transformed)
  712. log2fc = mean_group2 - mean_group1
  713. # t-test
  714. pvals = ttest_ind(group2.T, group1.T, axis=0, equal_var=False).pvalue
  715. # FDR correction (Benjamini-Hochberg)
  716. rejected, pvals_corrected, _, _ = multipletests(pvals, alpha=0.05, method='fdr_bh')
  717. # Create volcano dataframe
  718. volcano_df = pd.DataFrame({
  719. "Gene": df_expression.index,
  720. "log2FC": log2fc,
  721. "pval": pvals,
  722. "pval_adj": pvals_corrected,
  723. "significant": rejected
  724. })
  725. volcano_df["neg_log10_pval"] = -np.log10(volcano_df["pval_adj"].clip(lower=1e-300))
  726. volcano_df = volcano_df.replace([np.inf, -np.inf], np.nan).dropna()
  727. # Volcano plot then I
  728. plt.figure(figsize=(10, 6))
  729. plt.scatter(
  730. volcano_df["log2FC"],
  731. volcano_df["neg_log10_pval"],
  732. c=volcano_df["significant"].map({True: "red", False: "gray"}),
  733. alpha=0.7
  734. )
  735. # Label all genes with abs(log2FC) > 1 and FDR < 0.05
  736. highlighted = volcano_df[(volcano_df["significant"]) & (volcano_df["log2FC"].abs() > 1)]
  737. for _, row in highlighted.iterrows():
  738. plt.text(row["log2FC"], row["neg_log10_pval"], row["Gene"],
  739. fontsize=5, ha='center', va='bottom')
  740. # Threshold lines
  741. plt.axhline(-np.log10(0.05), linestyle='--', color='gray')
  742. plt.axvline(1, linestyle='--', color='gray')
  743. plt.axvline(-1, linestyle='--', color='gray')
  744. plt.title("Volcano Plot with FDR Correction")
  745. plt.xlabel("Log₂ Fold Change")
  746. plt.ylabel("-Log₁₀ Adjusted p-value")
  747. plt.tight_layout()
  748. plt.show()
  749. # %% [markdown]
  750. # ## log2fc = mean_group2 - mean_group1
  751. # ##So:
  752. #
  753. # ##If log2FC > 0, the gene is higher in Group 2
  754. #
  755. # ##If log2FC < 0, it’s higher in Group 1
  756. # %%
  757. sig_genes_df = volcano_df[
  758. (volcano_df["pval_adj"] < 0.05) &
  759. (volcano_df["log2FC"].abs() > 1)
  760. ]
  761. # Sort by most significant
  762. sig_genes_df = sig_genes_df.sort_values("pval_adj")
  763. # View top 10
  764. print(sig_genes_df.head(10))
  765. # Export all to Excel
  766. sig_genes_df.to_excel("significant_genes_filtered.xlsx", index=False)
  767. # %%
  768. # Calculate means
  769. mean_group1 = group1.mean(axis=1)
  770. mean_group2 = group2.mean(axis=1)
  771. # Log2 fold change (already log-transformed)
  772. log2fc = mean_group2 - mean_group1
  773. # t-test
  774. pvals = ttest_ind(group2.T, group1.T, axis=0, equal_var=False).pvalue
  775. # FDR correction (Benjamini-Hochberg)
  776. rejected, pvals_corrected, _, _ = multipletests(pvals, alpha=0.05, method='fdr_bh')
  777. # Create volcano dataframe
  778. volcano_df = pd.DataFrame({
  779. "Gene": df_expression.index,
  780. "log2FC": log2fc,
  781. "pval": pvals,
  782. "pval_adj": pvals_corrected,
  783. "significant": rejected
  784. })
  785. volcano_df["neg_log10_pval"] = -np.log10(volcano_df["pval_adj"].clip(lower=1e-300))
  786. volcano_df = volcano_df.replace([np.inf, -np.inf], np.nan).dropna()
  787. # Volcano plot
  788. plt.figure(figsize=(10, 6))
  789. plt.scatter(
  790. volcano_df["log2FC"],
  791. volcano_df["neg_log10_pval"],
  792. c=volcano_df["significant"].map({True: "lightblue", False: "gray"}),
  793. alpha=0.7
  794. )
  795. # 🎯 Annotate specific genes of interest
  796. genes_to_annotate = ["SOX2", "NES", "MIK67", "POU3F", "NOTCH1", "SLC6A9", "GFAP", "GRM2", "GRM5", "RGS10", "FOXP1", "CD44", "ALDH1L1", "SYT7", "GLUL", "SYT11", "SCGN", "SLC6A1", "SLC6A9" "NOTCH1", "DCX", "SDC1", "EGFR", "PHPR"] # ← customize this
  797. # 1. Not significant
  798. df_gray = volcano_df[~volcano_df["significant"]]
  799. # 2. Significant but not in your gene list
  800. df_lightblue = volcano_df[
  801. (volcano_df["significant"]) & (~volcano_df["Gene"].isin(genes_to_annotate))
  802. ]
  803. # 3. Highlighted genes (e.g., NES, TP53, etc.)
  804. df_pink = volcano_df[volcano_df["Gene"].isin(genes_to_annotate)]
  805. plt.figure(figsize=(10, 6))
  806. ax = plt.gca()
  807. # Plot gray first
  808. ax.scatter(df_gray["log2FC"], df_gray["neg_log10_pval"],
  809. color="gray", alpha=0.6, s=30, label="Not significant")
  810. # Plot lightblue next
  811. ax.scatter(df_lightblue["log2FC"], df_lightblue["neg_log10_pval"],
  812. color="lightblue", alpha=0.7, s=30, label="Significant")
  813. # Plot pink last (so it’s on top)
  814. ax.scatter(df_pink["log2FC"], df_pink["neg_log10_pval"],
  815. color="bisque", edgecolor="darkorange", linewidth=1.0,
  816. s=60, label="Highlighted Genes")
  817. for _, row in df_pink.iterrows():
  818. ax.text(row["log2FC"], row["neg_log10_pval"], row["Gene"],
  819. fontsize=10, ha='center', va='bottom', fontweight='bold', color='black')
  820. plt.axhline(-np.log10(0.05), linestyle='--', color='gray')
  821. plt.axvline(1, linestyle='--', color='gray')
  822. plt.axvline(-1, linestyle='--', color='gray')
  823. plt.xlabel("Log₂ Fold Change", fontsize=16)
  824. plt.ylabel("−Log₁₀ Adjusted p-value", fontsize=16)
  825. plt.title("Volcano Plot with Layered Highlighting", fontsize=16)
  826. #plt.legend(frameon=True)
  827. plt.tight_layout()
  828. #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\Volcano plot.svg", format="svg")
  829. plt.show()
  830. # %%
  831. # 1. Filter volcano_df for significant DE genes
  832. sig_genes_df = volcano_df[
  833. (volcano_df["pval_adj"] < 0.05) &
  834. (volcano_df["log2FC"].abs() > 1)
  835. ].copy()
  836. # 2. Sort by adjusted p-value
  837. sig_genes_df = sig_genes_df.sort_values("pval_adj")
  838. # 3. (Optional) View top 10
  839. print("Top significant genes:\n", sig_genes_df.head(10))
  840. # 4. Export to Excel
  841. output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\sig_genes.xlsx"
  842. #sig_genes_df.to_excel(output_path, index=False)
  843. print(f"Exported to: {output_path}")
  844. # %%
  845. #!pip install gprofiler-official
  846. # %%
  847. from gprofiler import GProfiler
  848. # Get your gene list
  849. gene_list = sig_genes_df["Gene"].tolist()
  850. # Create gProfiler object
  851. gp = GProfiler(return_dataframe=True)
  852. # Run GO/pathway enrichment
  853. results = gp.profile(
  854. organism="hsapiens",
  855. query=gene_list,
  856. user_threshold=0.05,
  857. sources=["GO:BP", "GO:MF", "GO:CC", "KEGG", "REAC", "WP"]
  858. )
  859. # Step 4: Check results
  860. print(f"Submitted {len(gene_list)} genes.")
  861. if results.empty:
  862. print("⚠️ No significant enrichment found. Try relaxing thresholds or using more genes.")
  863. else:
  864. print("✅ Enrichment completed. Columns available:", results.columns.tolist())
  865. # Only show if 'term_name' exists
  866. if "term_name" in results.columns:
  867. print(results[["term_name", "p_value", "source"]].head(10))
  868. else:
  869. print("⚠️ 'term_name' column not found in results. Full results:")
  870. print(results.head())
  871. # Save results
  872. output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\gprofiler_enrichment_results.xlsx"
  873. #results.to_excel(output_path, index=False)
  874. print(f"✅ Saved results to {output_path}")
  875. # %%
  876. print("Columns in results:", results.columns.tolist())
  877. print(results.head())
  878. # %%
  879. print(f"Number of enriched terms found: {results.shape[0]}")
  880. # %%
  881. import seaborn as sns
  882. import matplotlib.pyplot as plt
  883. # Make sure results aren't empty
  884. if not results.empty and "name" in results.columns:
  885. # Step 1: Pick top N terms (you can change 10)
  886. top_terms = results.sort_values("p_value").head(10)
  887. # Step 2: Barplot of top pathways
  888. plt.figure(figsize=(8, 6))
  889. sns.barplot(
  890. y="name",
  891. x="-log10(p_value)",
  892. data=top_terms.assign(**{"-log10(p_value)": -np.log10(top_terms["p_value"])}),
  893. palette="viridis"
  894. )
  895. plt.title("Top 10 Enriched Pathways (g:Profiler)", fontsize=16)
  896. plt.xlabel("-Log₁₀ p-value", fontsize=14)
  897. plt.ylabel("Pathway", fontsize=14)
  898. plt.tight_layout()
  899. plt.show()
  900. else:
  901. print("⚠️ No results to plot.")
  902. # %%
  903. # Find the row for the pathway you care about
  904. row = results[results["name"] == "cell junction"]
  905. if not row.empty:
  906. print("Matched Genes:", row["intersection_size"].values[0])
  907. else:
  908. print("⚠️ No pathway found with that exact name!")
  909. # Show the genes (usually in 'intersection')
  910. print(row["intersection_size"].values[0])
  911. # %%
  912. import seaborn as sns
  913. import matplotlib.pyplot as plt
  914. import numpy as np
  915. # Confirm 'results' is populated and has expected columns
  916. if not results.empty and "name" in results.columns and "p_value" in results.columns:
  917. # Step 1: Compute -log10(p-value)
  918. results["log_p"] = -np.log10(results["p_value"])
  919. # Step 2: Get top 10 terms by significance
  920. top_terms = results.nsmallest(10, "p_value")
  921. # Step 3: Bubble plot
  922. plt.figure(figsize=(10, 8))
  923. sns.scatterplot(
  924. data=top_terms,
  925. x="intersection_size", # Number of DEGs in each term
  926. y="name", # GO term or pathway name
  927. size="intersection_size",
  928. hue="log_p",
  929. sizes=(100, 1000),
  930. palette="viridis",
  931. legend="brief"
  932. )
  933. plt.xlabel("Gene Count in Term", fontsize=12)
  934. plt.ylabel("Pathway / GO Term", fontsize=12)
  935. plt.title("Top Enriched Terms (Bubble Plot)", fontsize=14)
  936. plt.tight_layout()
  937. plt.show()
  938. else:
  939. print("⚠️ No valid enrichment results found to plot.")
  940. # %%
  941. import matplotlib.pyplot as plt
  942. import seaborn as sns
  943. import numpy as np
  944. # List of GO terms or pathway names you're interested in (case-insensitive match)
  945. terms_of_interest = [
  946. "cell differentiation",
  947. "astrocyte differentiation",
  948. "nervous system development",
  949. "neurogenesis",
  950. "generation of neurons",
  951. "neuron differentiation",
  952. "cortical cytoskeleton organization",
  953. "synapse",
  954. "axon development",
  955. "gliogenesis",
  956. "Neuroinflammation and glutamatergic signaling",
  957. ]
  958. # --- 1. Prepare p-value for color scale ---
  959. filtered_terms["p_value_capped"] = filtered_terms["p_value"].clip(lower=1e-10) # avoid log(0)
  960. filtered_terms["p_label"] = filtered_terms["p_value"].apply(lambda x: f"{x:.1e}")
  961. # Sort to control y-axis order (optional)
  962. filtered_terms = filtered_terms.sort_values("intersection_size", ascending=True).reset_index(drop=True)
  963. # --- Calculate plot boundaries safely ---
  964. max_x = filtered_terms["intersection_size"].max()
  965. x_offset = max_x * 0.40 # shift p-values to the right by ~10%
  966. x_buffer = max_x * 0.25 # add buffer to x-axis so text isn't cut off
  967. # --- Set up plot ---
  968. fig, ax = plt.subplots(figsize=(12, 8))
  969. bubble = sns.scatterplot(
  970. data=filtered_terms,
  971. x="intersection_size",
  972. y="name",
  973. size="intersection_size",
  974. hue="p_value",
  975. palette="Blues_r",
  976. sizes=(200, 1200),
  977. alpha=0.8,
  978. edgecolor="gray",
  979. legend="brief",
  980. ax=ax
  981. )
  982. # --- Add p-value text BELOW bubble ---
  983. for i, row in filtered_terms.iterrows():
  984. ax.text(
  985. row["intersection_size"],
  986. i - 0.5, # further down
  987. row["p_label"],
  988. ha="left",
  989. va="top",
  990. fontsize=12,
  991. color="black"
  992. )
  993. # --- Adjust font sizes ---
  994. ax.set_yticks(range(len(filtered_terms)))
  995. ax.set_yticklabels(filtered_terms["name"], fontsize=14)
  996. ax.tick_params(axis='x', labelsize=12)
  997. ax.set_xlabel("Gene Count in Term", fontsize=14)
  998. ax.set_ylabel("GO Term", fontsize=14)
  999. ax.set_title("Selected GO Terms (Enrichment Bubble Plot)", fontsize=16)
  1000. # --- Prevent clipping of bubbles or labels ---
  1001. ax.set_xlim(0, max_x + x_buffer)
  1002. ax.set_ylim(-1, len(filtered_terms) - 0.5)
  1003. # --- Fix legend position outside plot ---
  1004. sns.move_legend(bubble, "upper left", bbox_to_anchor=(1.02, 1), frameon=True)
  1005. # --- Improve layout ---
  1006. plt.grid(axis="x", linestyle="--", alpha=0.3)
  1007. plt.tight_layout(rect=[0, 0, 1, 1]) # leave space on right for legend
  1008. #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\Enrichment Bubble Plot.svg", format="svg", bbox_inches="tight")
  1009. plt.show()
  1010. # %%
  1011. # Example: Plot a single protein across conditions
  1012. proteins_of_interest = ["NES", "MAP2", "GFAP"] # Replace with your protein
  1013. # Make sure your expression data is in long format
  1014. # Melt data for seaborn plotting
  1015. expr_long = df_imputed.loc[proteins_of_interest].T.reset_index().melt(id_vars="index", var_name="Protein", value_name="Expression")
  1016. expr_long.rename(columns={"index": "Sample"}, inplace=True)
  1017. # Add group info from sample names if encoded
  1018. expr_long["Group"] = expr_long["Sample"].str[0] # Adjust if needed
  1019. # Plot
  1020. plt.figure(figsize=(12, 6))
  1021. sns.boxplot(data=expr_long, x="Protein", y="Expression", hue="Group")
  1022. sns.stripplot(data=expr_long, x="Protein", y="Expression", hue="Group", dodge=True, color='black', size=4, alpha=0.5)
  1023. plt.title("Expression of Selected Proteins by Group")
  1024. plt.xticks(rotation=45)
  1025. plt.tight_layout()
  1026. plt.show()
  1027. # %% [markdown]
  1028. # # df_expression = df_imputed # ⬅️ This was your imputed expression matrix df_imputed = df_expression # If df_expression is already defined
  1029. # %%
  1030. df_imputed = df_expression # If df_expression is already defined
  1031. # %%
  1032. df
  1033. # %%
  1034. # Assuming your DataFrame is named 'df'
  1035. column_sums = df.sum(axis=0) # Calculate the sum of each column
  1036. # Divide each element by the corresponding column sum
  1037. df_normalized = df.div(column_sums, axis=1)
  1038. df_normalized
  1039. # %%
  1040. # Handle missing values (e.g., impute with mean or median), imputation may contain very small numbers
  1041. # Create a KNN imputer object
  1042. imputer = KNNImputer(n_neighbors=5) # Adjust n_neighbors as needed
  1043. # Impute missing values using KNN
  1044. df_imputed = imputer.fit_transform(df_normalized)
  1045. # Create a new DataFrame with imputed values
  1046. df_imputed = pd.DataFrame(df_imputed, columns=df.columns, index= df.index)
  1047. df_imputed*10e6
  1048. # %%
  1049. import pandas as pd
  1050. import numpy as np
  1051. import matplotlib.pyplot as plt
  1052. import seaborn as sns
  1053. from scipy.stats import mannwhitneyu
  1054. import os
  1055. import math
  1056. # --- Scale expression data ---
  1057. df_scaled = df_imputed * 1e7
  1058. # --- Assign groups from sample names ---
  1059. group1_labels = ['A', 'B', 'C', 'D']
  1060. group2_labels = ['E', 'F', 'G', 'H']
  1061. sample_to_group = df_scaled.columns.str[0]
  1062. group_map = sample_to_group.map(lambda x: 'Group 1' if x in group1_labels else ('Group 2' if x in group2_labels else 'Other'))
  1063. # --- Convert to long-form dataframe ---
  1064. df_long = df_scaled.T.copy()
  1065. df_long["Group"] = group_map.values
  1066. df_long = df_long.reset_index().melt(id_vars=["index", "Group"], var_name="Gene", value_name="Expression")
  1067. df_long.rename(columns={"index": "Sample"}, inplace=True)
  1068. # --- Genes of interest ---
  1069. genes_of_interest = ["STAT3", "ZEB2", "GFAP", "PLPP3", "SOX9", "NR3C1"] # add more as needed
  1070. n_genes = len(genes_of_interest)
  1071. # --- Grid dimensions (2 columns) ---
  1072. n_cols = 2
  1073. n_rows = math.ceil(n_genes / n_cols)
  1074. # --- Set up the figure ---
  1075. fig, axes = plt.subplots(n_rows, n_cols, figsize=(14, 5 * n_rows), squeeze=False)
  1076. # --- Plot each gene in grid ---
  1077. for idx, gene in enumerate(genes_of_interest):
  1078. row = idx // n_cols
  1079. col = idx % n_cols
  1080. ax = axes[row][col]
  1081. df_gene = df_long[df_long["Gene"] == gene]
  1082. group1 = df_gene[df_gene["Group"] == "Group 1"]["Expression"]
  1083. group2 = df_gene[df_gene["Group"] == "Group 2"]["Expression"]
  1084. # Mann–Whitney U test
  1085. u_stat, p_val = mannwhitneyu(group1, group2, alternative="two-sided")
  1086. sig = "***" if p_val < 0.001 else "**" if p_val < 0.01 else "*" if p_val < 0.05 else "ns"
  1087. # Plot
  1088. sns.violinplot(data=df_gene, x="Group", y="Expression", palette=["lightblue", "lightgray"], ax=ax, inner=None)
  1089. sns.stripplot(data=df_gene, x="Group", y="Expression", color="black", size=4, jitter=True, alpha=0.6, ax=ax)
  1090. # Significance annotation
  1091. y_max = df_gene["Expression"].max()
  1092. ax.set_ylim(top=y_max + 0.25 * y_max)
  1093. ax.plot([0, 1], [y_max + 0.05 * y_max] * 2, lw=1.5, c="black")
  1094. ax.text(0.5, y_max + 0.1 * y_max, sig, ha="center", va="center", fontsize=14)
  1095. # Style
  1096. ax.set_title(f"{gene} (p = {p_val:.2e}, {sig})", fontsize=13)
  1097. ax.set_xlabel("")
  1098. ax.set_ylabel("Expression", fontsize=12)
  1099. ax.tick_params(axis="both", labelsize=11)
  1100. # --- Hide unused subplots if total isn't a full grid ---
  1101. for i in range(n_genes, n_rows * n_cols):
  1102. fig.delaxes(axes[i // n_cols][i % n_cols])
  1103. # --- Final layout and export ---
  1104. plt.tight_layout()
  1105. output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\violin_plots\multi_gene_violin_grid.svg"
  1106. #plt.savefig(output_path, format="svg", bbox_inches="tight")
  1107. plt.show()
  1108. # %%
  1109. import numpy as np
  1110. import pandas as pd
  1111. import matplotlib.pyplot as plt
  1112. import seaborn as sns
  1113. from scipy.stats import mannwhitneyu
  1114. # --- List of genes to analyze ---
  1115. genes_of_interest = ["STAT3", "ZEB2", "GFAP", "SOX9", "NR3C1", "NOTCH1", "QKI", "SYN1", "DLG4", "SLC17A7", "SLC17A6", "CCNE1", "CCNA2", "DCX", "EGFR", "SDC1", "SOX2", "PRPH"] # customize as needed
  1116. # --- Setup for results ---
  1117. results = []
  1118. for gene in genes_of_interest:
  1119. group1 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['A', 'B', 'C', 'D'])]
  1120. group2 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['E', 'F', 'G', 'H'])]
  1121. # Compute log2 fold change
  1122. fc = group1.mean() / group2.mean()
  1123. log2fc = np.log2(fc)
  1124. # Mann–Whitney U test
  1125. u_stat, p_val = mannwhitneyu(group1, group2, alternative='two-sided')
  1126. # Significance label
  1127. if p_val < 0.001:
  1128. sig = '***'
  1129. elif p_val < 0.01:
  1130. sig = '**'
  1131. elif p_val < 0.05:
  1132. sig = '*'
  1133. else:
  1134. sig = 'ns'
  1135. # Collect
  1136. results.append({"Gene": gene, "log2FC": log2fc, "pval": p_val, "sig": sig})
  1137. # --- Convert to DataFrame ---
  1138. fc_df = pd.DataFrame(results)
  1139. # --- Plot ---
  1140. plt.figure(figsize=(10, 6))
  1141. sns.barplot(data=fc_df, x="Gene", y="log2FC", color="lightblue", edgecolor="gray")
  1142. # Add significance stars above bars
  1143. for i, row in fc_df.iterrows():
  1144. y = row["log2FC"]
  1145. offset = 0.05 * np.max(np.abs(fc_df["log2FC"]))
  1146. plt.text(i, y + np.sign(y) * offset, row["sig"], ha='center', va='bottom' if y >= 0 else 'top', fontsize=14)
  1147. # Style
  1148. plt.axhline(0, color="black", lw=1)
  1149. plt.ylabel("log₂ Fold Change (Group1 / Group2)", fontsize=13)
  1150. plt.xlabel("")
  1151. plt.title("Fold Change Between Groups with Expression Significance", fontsize=14)
  1152. plt.xticks(fontsize=12)
  1153. plt.yticks(fontsize=12)
  1154. plt.tight_layout()
  1155. # Expand y-limits so stars fit
  1156. y_max = fc_df["log2FC"].max()
  1157. y_min = fc_df["log2FC"].min()
  1158. y_range = y_max - y_min
  1159. plt.ylim(y_min - 0.2 * y_range, y_max + 0.2 * y_range)
  1160. output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\foldchange astrocytes differentiation.svg"
  1161. #plt.savefig(output_path, format="svg", bbox_inches="tight")
  1162. plt.show()
  1163. # %%
  1164. # Display sorted fold change table
  1165. print(fc_df.sort_values("log2FC", ascending=False))
  1166. # %%
  1167. import numpy as np
  1168. import pandas as pd
  1169. import matplotlib.pyplot as plt
  1170. import seaborn as sns
  1171. from scipy.stats import mannwhitneyu
  1172. # --- List of genes to analyze ---
  1173. #genes_of_interest = ["SLC6A9", "GRM2", "PSAT1", "SHMT1", "GLUL", "GRM5", "SLC1A3", "GRIK1"] # customize as needed
  1174. genes_of_interest = ["STAT3", "ZEB2", "GFAP", "SOX9", "NR3C1", "NOTCH1", "QKI"]
  1175. # --- Setup for results ---
  1176. results = []
  1177. for gene in genes_of_interest:
  1178. group1 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['A', 'B', 'C', 'D'])]
  1179. group2 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['E', 'F', 'G', 'H'])]
  1180. # Compute log2 fold change
  1181. fc = group2.mean() / group1.mean()
  1182. log2fc = np.log2(fc)
  1183. # Mann–Whitney U test
  1184. u_stat, p_val = mannwhitneyu(group1, group2, alternative='two-sided')
  1185. # Significance label
  1186. if p_val < 0.001:
  1187. sig = '***'
  1188. elif p_val < 0.01:
  1189. sig = '**'
  1190. elif p_val < 0.05:
  1191. sig = '*'
  1192. else:
  1193. sig = 'ns'
  1194. # Collect
  1195. results.append({"Gene": gene, "log2FC": log2fc, "pval": p_val, "sig": sig})
  1196. # --- Convert to DataFrame ---
  1197. fc_df = pd.DataFrame(results)
  1198. # --- Prepare plot ---
  1199. plt.figure(figsize=(10, 6))
  1200. # Sort genes by fold change for visual clarity
  1201. fc_df_sorted = fc_df.sort_values("log2FC", ascending=False).reset_index(drop=True)
  1202. # Plot dots (log2FC on y-axis)
  1203. sns.stripplot(
  1204. data=fc_df_sorted,
  1205. x="log2FC",
  1206. y="Gene",
  1207. size=10,
  1208. color="lightblue",
  1209. edgecolor="orange",
  1210. )
  1211. # Add vertical line at 0 (no change)
  1212. plt.axvline(0, color="black", lw=1)
  1213. # Add significance stars next to dots
  1214. x_range = fc_df_sorted["log2FC"].max() - fc_df_sorted["log2FC"].min()
  1215. x_offset = 0.03 * x_range # offset stars slightly to the right
  1216. for i, row in fc_df_sorted.iterrows():
  1217. x = row["log2FC"]
  1218. y = i
  1219. plt.text(
  1220. x + x_offset,
  1221. y,
  1222. row["sig"],
  1223. va="center",
  1224. ha="left",
  1225. fontsize=13,
  1226. color="black"
  1227. )
  1228. # Style
  1229. plt.xlabel("log₂ Fold Change (Group1 / Group2)", fontsize=13)
  1230. plt.ylabel("")
  1231. plt.title("Fold Change by Gene with Significance", fontsize=14)
  1232. plt.xticks(fontsize=12)
  1233. plt.yticks(fontsize=12)
  1234. plt.tight_layout()
  1235. plt.show()
  1236. # %%
  1237. import pandas as pd
  1238. import seaborn as sns
  1239. import matplotlib.pyplot as plt
  1240. from scipy.stats import ttest_ind
  1241. # --- Protein list to plot ---
  1242. proteins_of_interest = ["STAT3", "ZEB2", "GFAP", "PLPP3", "SOX9", "NR3C1"] # Update as needed
  1243. # --- Melt expression data ---
  1244. expr_long = df_imputed.loc[proteins_of_interest].T.reset_index().melt(
  1245. id_vars="index", var_name="Protein", value_name="Expression"
  1246. )
  1247. expr_long.rename(columns={"index": "Sample"}, inplace=True)
  1248. # --- Assign groups using first letter of sample name ---
  1249. expr_long["Group"] = expr_long["Sample"].str[0]
  1250. # --- Plot: One violin plot per protein ---
  1251. plt.figure(figsize=(12, 6 * len(proteins_of_interest)))
  1252. for i, protein in enumerate(proteins_of_interest):
  1253. plt.subplot(len(proteins_of_interest), 1, i + 1)
  1254. # Subset data for this protein
  1255. data = expr_long[expr_long["Protein"] == protein]
  1256. # Group split
  1257. group_values = [g["Expression"].values for _, g in data.groupby("Group")]
  1258. if len(group_values) == 2:
  1259. stat, pval = ttest_ind(*group_values)
  1260. else:
  1261. pval = float('nan')
  1262. # Determine significance label
  1263. if pval < 0.001:
  1264. sig_label = "***"
  1265. elif pval < 0.01:
  1266. sig_label = "**"
  1267. elif pval < 0.05:
  1268. sig_label = "*"
  1269. else:
  1270. sig_label = "ns"
  1271. # Plot violin and strip
  1272. sns.violinplot(data=data, x="Group", y="Expression", inner=None, palette="pastel")
  1273. sns.stripplot(data=data, x="Group", y="Expression", color="black", size=4, alpha=0.6)
  1274. # Annotate significance
  1275. y_max = data["Expression"].max()
  1276. y_min = data["Expression"].min()
  1277. y_range = y_max - y_min
  1278. plt.text(0.5, y_max + 0.05 * y_range, sig_label, ha="center", fontsize=14)
  1279. # Titles and labels
  1280. plt.title(f"{protein} (p = {pval:.2e})")
  1281. plt.xlabel("Condition")
  1282. plt.ylabel("Expression")
  1283. plt.tight_layout()
  1284. plt.show()
  1285. # %%
  1286. # %%
  1287. # %%
  1288. # %%

PCA and Volcano plot.ipynb at commit 9be7e24, no license · at the source

Overview

Authors: Jesús Muñoz-Estrada1, Andrew Mostafania1, Lahiruni Halwatura1, Ali Haghani1, Yuming Jiang1, Jesse G Meyer1
ORCID iDs: Jesse G Meyer
  1. Department of Computational Biomedicine, Smidt Heart Institute, Cedars Sinai Medical Center, Los Angeles, California, USA
Institutions: Cedars-Sinai Medical Center (United States); Cedars-Sinai Smidt Heart Institute (United States)
Journal: Molecular & cellular proteomics : MCP, volume 25, issue 7, article 101604
Dates: received 7 July 2025; published online 17 June 2026; in print July 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1016/j.mcpro.2026.101604 · PMID 42309337 · PMCID PMC13382583 · OpenAlex W7164999840
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), cellular / molecular (subfield)
Methods: Preprocessing, Statistics, Smoothing, state filtering, decompositions, Connectivity, Machine learning
Keywords: KOLF2.1J, induced pluripotent stem cells (iPSCs), Neurogenin-2 (NGN2) stable integration, AAVS1, EF1α transgene silencing, single-cell proteomics, bulk proteomics, iPSC-derived neuronal-enriched cultures, Notch signaling, NGN2 dosage
MeSH: Basic Helix-Loop-Helix Proteins*, Gene Dosage*, Induced Pluripotent Stem Cells*, Nerve Tissue Proteins*, Neurons*, Animals, Cell Differentiation, Cells, Cultured, DNA Methylation, Humans, Promoter Regions, Genetic, Proteomics, Receptors, Notch (* major topic)
Topic: Peptidase Inhibition and Analysis (Oncology, Medicine), according to OpenAlex
Funding: NIA NIH HHS (R21 AG074234); National Institute on Aging; NIGMS NIH HHS (R35 GM142502); National Institute of General Medical Sciences
Citations: not cited yet (Europe PMC); 37 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

xomicsdatascience/generation_of_iNeurons

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 9be7e24aeb13dcdd0470bc641b06cf1d171ab335, 30 January 2026
Languages: Jupyter (10)
Size: 18 files, 10 scripts
Software Heritage: not archived
Found in: “Code Availability”
Holds: README, 10 notebooks
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (10 files), pandas (10 files), Matplotlib (9 files), seaborn (9 files), SciPy (6 files), scikit-learn (4 files), statsmodels (2 files), anndata (1 file), Plotly (1 file), Scanpy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
11 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1016/j.mcpro.2026.101604.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 10 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • no repository, dataset or request procedure was recognized in it

Read it in the paper: doi.org/10.1016/j.mcpro.2026.101604.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Authors: added Jesse G Meyer (0000-0003-2753-3926); removed Jesse G Meyer

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 10 keywords, 13 MeSH terms, 4 funders, 37 references.

Cite

This paper

Muñoz-Estrada, J., Mostafania, A., Halwatura, L., Haghani, A., Jiang, Y., & Meyer, J. G. (2026). Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures. Molecular & cellular proteomics : MCP, 25(7), 101604. https://doi.org/10.1016/j.mcpro.2026.101604

BibTeX

@article{munozestrada2026optimizing,
author = {Muñoz-Estrada, Jesús and Mostafania, Andrew and Halwatura, Lahiruni and Haghani, Ali and Jiang, Yuming and Meyer, Jesse G},
title = {{Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures}},
journal = {Molecular \& cellular proteomics : MCP},
year = {2026},
month = jun,
volume = {25},
number = {7},
pages = {101604},
publisher = {American Society for Biochemistry and Molecular Biology},
issn = {1535-9476},
doi = {10.1016/j.mcpro.2026.101604},
url = {https://doi.org/10.1016/j.mcpro.2026.101604},
pmid = {42309337},
pmcid = {PMC13382583}
}

RIS

TY - JOUR
AU - Muñoz-Estrada, Jesús
AU - Mostafania, Andrew
AU - Halwatura, Lahiruni
AU - Haghani, Ali
AU - Jiang, Yuming
AU - Meyer, Jesse G
TI - Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures
T2 - Molecular & cellular proteomics : MCP
J2 - Mol Cell Proteomics
PY - 2026
DA - 2026/06/17
VL - 25
IS - 7
SP - 101604
SN - 1535-9476
PB - American Society for Biochemistry and Molecular Biology
DO - 10.1016/j.mcpro.2026.101604
UR - https://doi.org/10.1016/j.mcpro.2026.101604
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.mcpro.2026.101604",
"type": "article-journal",
"title": "Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures",
"container-title": "Molecular & cellular proteomics : MCP",
"author": [
{
"family": "Muñoz-Estrada",
"given": "Jesús"
},
{
"family": "Mostafania",
"given": "Andrew"
},
{
"family": "Halwatura",
"given": "Lahiruni"
},
{
"family": "Haghani",
"given": "Ali"
},
{
"family": "Jiang",
"given": "Yuming"
},
{
"family": "Meyer",
"given": "Jesse G"
}
],
"container-title-short": "Mol Cell Proteomics",
"volume": "25",
"issue": "7",
"page": "101604",
"DOI": "10.1016/j.mcpro.2026.101604",
"PMID": "42309337",
"PMCID": "PMC13382583",
"ISSN": "1535-9476",
"publisher": "American Society for Biochemistry and Molecular Biology",
"URL": "https://doi.org/10.1016/j.mcpro.2026.101604",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
17
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/nar/gkag368 [code]
Single-cell trajectory inference for detecting transient events in biological processes.
Journal: Nucleic acids research
In common: anndata, Scanpy, scikit-learn, 4 other tools, genetics / omics, cellular / molecular, 3 references, author Jesse G Meyer
[2] doi:10.1038/s41467-026-71803-3 [code]
Charting the transition from in vitro gliogenesis to the in vivo maturation of human glial progenitor cells transplanted into the hypomyelinated mouse brain.
Journal: Nature communications
In common: anndata, Scanpy, Plotly, 6 other tools, genetics / omics, cellular / molecular, 2 references
[3] doi:10.1038/s41514-026-00391-9 [code]
Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.
Journal: npj aging
In common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics, cellular / molecular, 1 reference
[4] doi:10.1038/s41593-026-02267-3 [code]
Spatial proteomic analysis in human Alzheimer's disease brains enables identification of microenvironment-dependent microglial cell states.
Journal: Nature neuroscience
In common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics, cellular / molecular
[5] doi:10.1126/sciadv.aeg3223 [code]
The extreme diversity of retinal amacrine cells has deep evolutionary roots.
Journal: Science advances
In common: anndata, Scanpy, Plotly, 6 other tools, genetics / omics, cellular / molecular, 1 reference
[6] doi:10.1371/journal.pcbi.1014346 [code]
StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion.
Journal: PLoS computational biology
In common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics
[7] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics
[8] doi:10.1016/j.celrep.2026.117270 [code]
Blocking apoptosis promotes survival and alters developmental dynamics of human retinal ganglion cells in retinal organoids.
Journal: Cell reports
In common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics
[9] doi:10.3390/ijms27167275 [code]
Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling.
Journal: International journal of molecular sciences
In common: anndata, Scanpy, statsmodels, 6 other tools, genetics / omics, cellular / molecular, 1 reference
[10] doi:10.1038/s41467-026-76054-w [code]
Distinct differentiation trajectories leave lasting impacts on gene regulation and function of V2a neurons.
Journal: Nature communications
In common: Scanpy, Plotly, pandas, 2 other tools, cellular / molecular, 4 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.