Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures.
The 7 matches
- [1] § Experimental Procedures › Bioinformatic and Computational Analysis ↔ python/PCA and Volcano plot.ipynb, lines 1063–1097 · score 0.79 · gProfiler, CC, KEGG, MF, REAC, WP
- [2] § Experimental Procedures › Bioinformatic and Computational Analysis ↔ python/2025-ipscs-scp-github.ipynb, lines 220–223 · score 0.61 · UMAP, Scanpy, pp, resolution, neighbors, Leiden
- [3] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/Cluster heatmap and centroid distances.ipynb, lines 298–330 · score 0.60 · highlights small distances, pairwise Euclidean distances, clustering heatmap, proteomic
- [4] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/PCA and Volcano plot.ipynb, lines 718–744 · score 0.59 · ALDH1L1, CD44, MKI67, TUBB3, SOX2, SYN1
- [5] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/PCA and Volcano plot.ipynb, lines 1177–1261 · score 0.56 · GO term, gliogenesis, glutamatergic, astrocytes, volcano, enrichment
- [6] § Results › Proteomic Analysis and Molecular Characterization of NGN2-Induced iPSC Transgenic Cultures ↔ python/PCA and Volcano plot.ipynb, lines 1389–1457 · score 0.55 · Log2 fold change, Mann Whitney, NOTCH1, Volcano, bar, astrocytes
- [7] § Experimental Procedures › Linear Modeling ↔ python/Baseline linear model .ipynb, lines 120–156 · score 0.52 · linear model, OLS, ANOVA, fitted, row, Log2
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 1,620 lines · 49 KB · no license · 4 matches
- # %%
- import pandas as pd
- import plotly.express as px
- from sklearn.decomposition import PCA
- from sklearn.impute import KNNImputer
- import numpy as np
- import matplotlib.pyplot as plt
- import pandas as pd
- import seaborn as sns
- import matplotlib.pyplot as plt
- from scipy.spatial.distance import pdist, squareform
- # %%
- import os
- os.getcwd()
- # %%
- # Load your spreadsheet data into a Pandas DataFrame
- df = pd.read_csv(r"..\\data\2025\dox-mzml.pg_matrix (1).tsv", sep='\t', index_col=2)
- df
- df.drop(['Protein.Group', 'Protein.Names', 'First.Protein.Description','N.Sequences','N.Proteotypic.Sequences'],axis=1, inplace=True)
- df
- # %%
- df.columns = [col.split("\\")[-1].split(".")[0] for col in df.columns]
- df
- # %%
- np.log(df).boxplot(fontsize=5)
- #plt.axhline(df.count(axis=0), color='red', linestyle='--', label='Mean')
- plt.title('raw log2 quantities per sample')
- #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment/boxquant_raw.svg", bbox_inches='tight')
- plt.show()
- # %%
- df.count(axis=0)
- # %%
- plt.rcParams['figure.figsize'] = 3,3
- #plt.rcParams['axes.grid'] = False
- df.count(axis=0).hist(bins=20, grid=False)
- plt.axvline(x=7500, color='black', linewidth=3)
- plt.xlabel('number of protein IDs')
- plt.ylabel('count')
- #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment/proteins_per_condition.svg", bbox_inches='tight')
- plt.show()
- # %%
- df.T.index
- # %%
- dropthese = df.T.index[df.count()<7500]
- dropthese
- # %%
- df = df.drop(dropthese, axis=1)
- # %%
- df.count()
- # %%
- # Assuming your DataFrame is named 'df'
- column_sums = df.sum(axis=0) # Calculate the sum of each column
- # Divide each element by the corresponding column sum
- df_normalized = df.div(column_sums, axis=1)
- df_normalized
- # %%
- # Handle missing values (e.g., impute with mean or median), imputation may contain very small numbers
- # Create a KNN imputer object
- imputer = KNNImputer(n_neighbors=5) # Adjust n_neighbors as needed
- # Impute missing values using KNN
- df_imputed = imputer.fit_transform(df_normalized)
- # Create a new DataFrame with imputed values
- df_imputed = pd.DataFrame(df_imputed, columns=df.columns, index= df.index)
- df_imputed*10e6
- # %%
- #[0,:] all columns in the first row, what type of distribution
- df_imputed.iloc[0,:].hist()
- # Set the x and y labels
- plt.xlabel('Value') # Label for x-axis
- plt.ylabel('Frequency') # Label for y-axis
- plt.show()
- # %%
- df_imputed = np.log(df_imputed*10e6) #np.log (...) applies the natural algorithm (log base e) to all values, improve linearity for PCA
- #make data more normally distributed and reduce the effects of outliers
- # %%
- df_imputed.iloc[0,:].hist()
- plt.xlabel('Value') # Label for x-axis
- plt.ylabel('Frequency') # Label for y-axis
- plt.show()
- # %%
- # Extract group labels from column names
- # Jessi code=group_labels = [int(col[-1]) for col in df.columns]
- #My group_labels are in the firts row
- group_labels = [col[0] for col in df.columns]
- group_labels
- # %%
- #shape of the transposed values
- df_imputed.values.T.shape
- # %%
- # Transpose the DataFrame
- df_transposed = df_imputed.T #T stands for transpose it flips the data frame rows into columns and columns into rows
- #samples as rows are requiered for PCA and StandardScaler
- # %%
- # Print the new shape
- print(df_transposed.shape)
- # %%
- df_transposed
- # %%
- # Standardize features (per column)
- from sklearn.preprocessing import StandardScaler
- scaler = StandardScaler()
- df_scaled = scaler.fit_transform(df_imputed.values) #df_scaled standarized
- # %%
- # Reconstruct a DataFrame from df_scaled for easier plotting
- df_scaled_df = pd.DataFrame(df_scaled, index=df_imputed.index, columns=df_imputed.columns)
- # Boxplot per sample (i.e., each box is one sample)
- plt.figure(figsize=(14, 6))
- df_scaled_df.boxplot(showfliers=False) # Set to True if you want to show outliers
- # Plot formatting
- plt.title("Boxplot of Standardized Feature Values per Sample", fontsize=16)
- plt.xlabel("Samples", fontsize=14)
- plt.ylabel("Standardized Feature Values", fontsize=14)
- plt.xticks(rotation=90, fontsize=10)
- plt.yticks(fontsize=12)
- plt.tight_layout()
- plt.show()
- # %%
- print("Shape of df_scaled_df:", df_scaled_df.shape)
- # %%
- print("Any missing values?", df_scaled_df.isnull().values.any())
- # %%
- print("Overall mean:", df_scaled_df.values.mean())
- print("Overall std:", df_scaled_df.values.std())
- # %%
- df_scaled_T = df_scaled_df.T # Now shape is (37, 9782)
- # %%
- print("Shape of df_scaled_df:", df_scaled_df.T.shape)
- # %%
- print("Any missing values?", df_scaled_df.T.isnull().values.any())
- # %%
- print("Overall mean:", df_scaled_df.T.values.mean())
- print("Overall std:", df_scaled_df.T.values.std())
- # %%
- # Transpose so samples are rows
- df_scaled_T = df_scaled_df.T
- # ✅ Convert all column names (features) to strings
- df_scaled_T.columns = df_scaled_T.columns.astype(str)
- # PCA
- from sklearn.decomposition import PCA
- pca = PCA(n_components=2)
- pca_data = pca.fit_transform(df_scaled_T)
- # Step 2: Create DataFrame with PCA results
- pca_df = pd.DataFrame(pca_data, columns=["PC1", "PC2"])
- pca_df["Sample"] = df_scaled_T.index # Sample names from rows
- pca_df["Group"] = pca_df["Sample"].str.split("_").str[0] # Customize based on your naming # not good for
- # Step 3: Plot with Matplotlib
- fig, ax = plt.subplots(figsize=(10, 8))
- # Color by group
- for group in pca_df["Group"].unique():
- group_data = pca_df[pca_df["Group"] == group]
- ax.scatter(group_data["PC1"], group_data["PC2"], label=f"{group}", s=150, alpha=0.7)
- # Optional: annotate each sample
- for _, row in pca_df.iterrows():
- ax.text(row["PC1"] + 0.3, row["PC2"] + 0.3, row["Sample"], fontsize=9, alpha=0.8)
- # Variance explained for axis labels
- explained_var = PCA(n_components=2).fit(df_scaled_T).explained_variance_ratio_ * 100
- ax.set_xlabel(f"PC1 ({explained_var[0]:.1f}%)", fontsize=14)
- ax.set_ylabel(f"PC2 ({explained_var[1]:.1f}%)", fontsize=14)
- # Final formatting
- ax.set_title("PCA Plot (Matplotlib)", fontsize=16, fontweight="bold")
- ax.legend(title="Group", fontsize=12)
- ax.tick_params(labelsize=12)
- plt.tight_layout()
- plt.show()
- # %%
- # Transpose so samples are rows
- df_scaled_T = df_scaled_df.T
- # ✅ Convert all column names (features) to strings
- df_scaled_T.columns = df_scaled_T.columns.astype(str)
- # PCA
- from sklearn.decomposition import PCA
- pca = PCA(n_components=2)
- pca_data = pca.fit_transform(df_scaled_T)
- # Step 2: Create DataFrame with PCA results
- pca_df = pd.DataFrame(pca_data, columns=["PC1", "PC2"])
- pca_df["Sample"] = df_scaled_T.index # Sample names from rows
- pca_df["Group"] = pca_df["Sample"].str[0] # Customize based on your naming # not good for
- # Step 3: Plot with Matplotlib
- fig, ax = plt.subplots(figsize=(12, 10))
- # Color by group
- for group in pca_df["Group"].unique():
- group_data = pca_df[pca_df["Group"] == group]
- ax.scatter(group_data["PC1"], group_data["PC2"], label=f"{group}", s=300, alpha=0.7) #dot size s=400
- # Optional: annotate each sample
- for _, row in pca_df.iterrows():
- ax.text(row["PC1"] + 0.3, row["PC2"] + 0.3, row["Sample"], fontsize=9, alpha=0.8)
- # Variance explained for axis labels
- explained_var = PCA(n_components=2).fit(df_scaled_T).explained_variance_ratio_ * 100
- ax.set_xlabel(f"PC1 ({explained_var[0]:.1f}%)", fontsize=14)
- ax.set_ylabel(f"PC2 ({explained_var[1]:.1f}%)", fontsize=14)
- # Final formatting
- ax.set_title("PCA Plot (Matplotlib)", fontsize=16, fontweight="bold") #axis label fontsize+
- ax.legend(title="Group", fontsize=12)
- ax.tick_params(labelsize=14)
- plt.tight_layout()
- plt.show()
- # %%
- # Step 1: Perform PCA
- pca = PCA(n_components=2)
- pca_data = pca.fit_transform(df_scaled_T)
- # Step 2: Create DataFrame with PCA results
- pca_df = pd.DataFrame(pca_data, columns=["PC1", "PC2"])
- pca_df["Sample"] = df_scaled_T.index
- pca_df["Group"] = pca_df["Sample"].str[0] # Adjust if needed
- # Step 3: Plot
- fig, ax = plt.subplots(figsize=(12, 10))
- # Plot each group with larger dots
- for group in pca_df["Group"].unique():
- group_data = pca_df[pca_df["Group"] == group]
- ax.scatter(
- group_data["PC1"],
- group_data["PC2"],
- label=f"Group {group}",
- s=700,
- alpha=0.7
- )
- # Axis labels with % variance explained
- explained = pca.explained_variance_ratio_ * 100
- #ax.set_title("PCA Plot", fontsize=20, fontweight='bold')
- ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=16)
- ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=16)
- # Make axis lines (spines) thicker
- ax.spines['bottom'].set_linewidth(2)
- ax.spines['left'].set_linewidth(2)
- ax.spines['top'].set_linewidth(2)
- ax.spines['right'].set_linewidth(2)
- # Bigger tick labels
- ax.tick_params(width=2, length=6, labelsize=25)
- # Legend
- ax.legend(title="Group", fontsize=12, title_fontsize=14, bbox_to_anchor=(1.05, 1), loc='upper left')
- # Tight layout
- plt.tight_layout()
- # Step 4: Export as SVG
- #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\PCA_final.svg", format="svg")
- # Optional: Show plot
- plt.show()
- # %%
- from sklearn.decomposition import PCA
- pca = PCA(n_components=2)
- pca_data = pca.fit_transform(df_scaled_T)
- explained = pca.explained_variance_ratio_ * 100
- print("Explained variance:", pca.explained_variance_ratio_)
- # %%
- print("Groups detected:", pca_df["Group"].unique())
- print("Number of groups:", pca_df["Group"].nunique())
- # %%
- from sklearn.cluster import KMeans
- # Step 1: Cluster PCA data (2D) into 2 clusters
- kmeans = KMeans(n_clusters=2, random_state=42)
- pca_df["Cluster"] = kmeans.fit_predict(pca_df[["PC1", "PC2"]])
- # Step 2: Plot with cluster colors instead of groups
- fig, ax = plt.subplots(figsize=(10, 8))
- # Define cluster colors
- colors = ['red', 'blue'] # Or use more for more clusters
- for cluster in pca_df["Cluster"].unique():
- cluster_data = pca_df[pca_df["Cluster"] == cluster]
- ax.scatter(cluster_data["PC1"], cluster_data["PC2"],
- label=f"Cluster {cluster+1}", c=colors[cluster],
- s=200, alpha=0.7)
- # Axes and legend
- explained = pca.explained_variance_ratio_ * 100
- ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
- ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
- ax.set_title("PCA with KMeans Clustering", fontsize=16, fontweight="bold")
- ax.legend(title="KMeans Cluster", fontsize=12)
- ax.tick_params(labelsize=12)
- # Optional: Thicker axes
- for spine in ax.spines.values():
- spine.set_linewidth(2)
- ax.tick_params(width=2, length=6)
- plt.tight_layout()
- plt.show()
- # %%
- pd.crosstab(pca_df["Group"], pca_df["Cluster"])
- # %%
- from sklearn.cluster import KMeans
- import matplotlib.pyplot as plt
- # Step 1: Cluster PCA data (2 clusters for example)
- kmeans = KMeans(n_clusters=2, random_state=42)
- pca_df["Cluster"] = kmeans.fit_predict(pca_df[["PC1", "PC2"]])
- # Step 2: Get centroids from the KMeans model
- centroids = kmeans.cluster_centers_ # shape: (n_clusters, 2)
- # Step 3: Plot
- fig, ax = plt.subplots(figsize=(10, 8))
- # Cluster colors
- colors = ['lightblue', 'gray']
- for cluster in pca_df["Cluster"].unique():
- cluster_data = pca_df[pca_df["Cluster"] == cluster]
- ax.scatter(
- cluster_data["PC1"],
- cluster_data["PC2"],
- label=f"Cluster {cluster+1}",
- c=colors[cluster],
- s=500,
- alpha=0.7
- )
- # Step 4: Plot centroids
- ax.scatter(
- centroids[:, 0], # PC1
- centroids[:, 1], # PC2
- marker='X',
- s=300,
- c='orange',
- label='Centroid',
- edgecolor='white',
- linewidth=2
- )
- # Axis formatting
- explained = pca.explained_variance_ratio_ * 100
- ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=22)
- ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=22)
- ax.set_title("PCA with KMeans Clustering and Centroids", fontsize=18, fontweight="bold")
- ax.tick_params(labelsize=22)
- for spine in ax.spines.values():
- spine.set_linewidth(2)
- ax.tick_params(width=2, length=6)
- #ax.legend(title="Cluster", fontsize=16)
- plt.tight_layout()
- #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\PCA_withKMeans.svg", format="svg")
- plt.show()
- # %%
- import seaborn as sns
- import matplotlib.pyplot as plt
- # Step 1: Transpose to get samples as rows
- df_plot = df_imputed.T
- # Step 2: Add group info based on sample names
- df_plot["Group"] = df_plot.index.str[0]
- # Step 3: Choose gene of interest (make sure it's in the column names now)
- gene_of_interest = "PRPH" # Replace with a valid gene name from df_imputed.index
- # Step 4: Move the gene from index to column if necessary
- df_plot.columns = df_plot.columns.astype(str)
- if gene_of_interest not in df_plot.columns:
- df_plot[gene_of_interest] = df_imputed.loc[gene_of_interest].values
- # Step 5: Violin plot
- plt.figure(figsize=(10, 6))
- sns.violinplot(data=df_plot, x="Group", y=gene_of_interest, inner="box", palette="Set2")
- # Step 6: Plot styling
- plt.title(f"Expression of {gene_of_interest} Across Groups", fontsize=16)
- plt.xlabel("Group", fontsize=14)
- plt.ylabel("Expression Level", fontsize=14)
- plt.xticks(fontsize=12)
- plt.yticks(fontsize=12)
- plt.tight_layout()
- plt.show()
- # %%
- print(df_imputed.index[:10].tolist())
- # %%
- import seaborn as sns
- import matplotlib.pyplot as plt
- # Step 1: Transpose to get samples as rows
- df_plot = df_imputed.T
- # Step 2: Add group info based on sample names
- df_plot["Group"] = df_plot.index.str[0]
- # Step 3: Choose gene of interest (make sure it's in the column names now)
- gene_of_interest = "SYN1" # Replace with a valid gene name from df_imputed.index
- # Step 4: Move the gene from index to column if necessary
- df_plot.columns = df_plot.columns.astype(str)
- if gene_of_interest not in df_plot.columns:
- df_plot[gene_of_interest] = df_imputed.loc[gene_of_interest].values
- # Step 5: Violin plot
- plt.figure(figsize=(10, 6))
- sns.violinplot(data=df_plot, x="Group", y=gene_of_interest, inner="box", palette="Set2")
- # Step 6: Plot styling
- plt.title(f"Expression of {gene_of_interest} Across Groups", fontsize=16)
- plt.xlabel("Group", fontsize=14)
- plt.ylabel("Expression Level", fontsize=14)
- plt.xticks(fontsize=12)
- plt.yticks(fontsize=12)
- plt.tight_layout()
- plt.show()
- # %%
- import pandas as pd
- # Step 1: Transpose to get samples as rows
- df_plot = df_imputed.T.copy() # Now shape: (samples, genes)
- # Step 2: Add group info (based on sample name, e.g., 'A1' → group 'A')
- df_plot["Group"] = df_plot.index.str[0] # Adjust if your sample names are different
- # Step 3: Extract expression of the gene of interest (e.g., TP53)
- gene_of_interest = "SLC17A6" # Make sure it's a valid gene in df_imputed.index
- gene_data = df_plot[[gene_of_interest, "Group"]] # Keep only gene and group
- # Step 4: Export to Excel
- output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\SC17A6_expression_by_group.xlsx"
- #gene_data.to_excel(output_path)
- print("Exported:", output_path)
- # %% [markdown]
- # ### Differential expression between clusters.
- # %%
- # Step 1: Transpose df_imputed to get samples as rows
- df_expr = df_imputed.T.copy() # shape: (samples, proteins)
- # Step 2: Add Cluster labels to expression matrix
- df_expr["Cluster"] = pca_df.set_index("Sample").loc[df_expr.index, "Cluster"]
- # Step 3: Group by cluster and compute mean expression
- cluster_means = df_expr.groupby("Cluster").mean().T # shape: (proteins, 2 clusters)
- # Step 4: Compute absolute difference between clusters
- cluster_means["diff"] = abs(cluster_means[0] - cluster_means[1])
- # Step 5: Sort by difference and get top 10 proteins
- top_proteins = cluster_means.sort_values("diff", ascending=False).head(20)
- print("Top 20 proteins separating clusters:\n")
- print(top_proteins)
- # %%
- # Cross-tabulate clusters and original groups
- cluster_group_counts = pd.crosstab(pca_df["Cluster"], pca_df["Group"])
- # Display the mapping
- print(cluster_group_counts)
- # %%
- cluster_group_counts.T.plot(kind="bar", stacked=True, figsize=(10, 6), colormap="Set2")
- plt.title("Group Composition per Cluster")
- plt.xlabel("Group")
- plt.ylabel("Sample Count")
- plt.legend(title="KMeans Cluster")
- plt.tight_layout()
- plt.show()
- # %%
- import matplotlib.pyplot as plt
- import seaborn as sns
- import pandas as pd
- # Setup colors and marker shapes
- group_colors = {
- "A": "forestgreen", "B": "red", "C": "dodgerblue", "D": "darkorange",
- "E": "gold", "F": "tomato", "G": "purple", "H": "skyblue"
- }
- cluster_markers = {
- 0: "o", # circle
- 1: "s", # square
- }
- # Create figure
- fig, ax = plt.subplots(figsize=(12, 10))
- # Plot each sample by group (color) and cluster (shape)
- for cluster in sorted(pca_df["Cluster"].unique()):
- for group in sorted(pca_df["Group"].unique()):
- subset = pca_df[(pca_df["Cluster"] == cluster) & (pca_df["Group"] == group)]
- ax.scatter(
- subset["PC1"],
- subset["PC2"],
- color=group_colors.get(group, "gray"),
- marker=cluster_markers.get(cluster, "o"),
- s=200,
- alpha=0.8,
- label=f"Group {group}, Cluster {cluster}"
- )
- # Axis formatting
- explained = pca.explained_variance_ratio_ * 100
- ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
- ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
- ax.set_title("PCA: Color by Group, Shape by Cluster", fontsize=18, fontweight="bold")
- ax.tick_params(labelsize=12)
- # Legend (no duplicates)
- handles, labels = ax.get_legend_handles_labels()
- unique_labels = dict(zip(labels, handles))
- ax.legend(unique_labels.values(), unique_labels.keys(), title="Group + Cluster", bbox_to_anchor=(1.05, 1), loc='upper left')
- plt.tight_layout()
- plt.show()
- # %%
- import matplotlib.patches as mpatches
- import numpy as np
- fig, ax = plt.subplots(figsize=(12, 10))
- # Plot each group with unique color, marker shape based on cluster
- for cluster in sorted(pca_df["Cluster"].unique()):
- for group in sorted(pca_df["Group"].unique()):
- subset = pca_df[(pca_df["Group"] == group) & (pca_df["Cluster"] == cluster)]
- if subset.empty:
- continue
- # Plot points
- ax.scatter(
- subset["PC1"],
- subset["PC2"],
- color=group_colors.get(group, "gray"),
- marker=cluster_markers.get(cluster, "o"),
- s=200,
- alpha=0.7,
- label=f"{group}, Cluster {cluster}"
- )
- # Add ellipse around the group
- if len(subset) >= 3: # Need at least 3 points to define an ellipse
- cov = np.cov(subset[["PC1", "PC2"]].T)
- lambda_, v = np.linalg.eig(cov)
- lambda_ = np.sqrt(lambda_)
- ell = mpatches.Ellipse(
- xy=(subset["PC1"].mean(), subset["PC2"].mean()),
- width=lambda_[0] * 4, # 2 std dev
- height=lambda_[1] * 4,
- angle=np.rad2deg(np.arccos(v[0, 0])),
- edgecolor=group_colors.get(group, "gray"),
- facecolor='none',
- linewidth=2,
- linestyle="--",
- alpha=0.7
- )
- ax.add_patch(ell)
- # Axes
- explained = pca.explained_variance_ratio_ * 100
- ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
- ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
- ax.set_title("PCA with Ellipses by Group and Clustering Shape", fontsize=18)
- # Legend cleanup
- handles, labels = ax.get_legend_handles_labels()
- unique = dict(zip(labels, handles))
- ax.legend(unique.values(), unique.keys(), title="Group + Cluster", bbox_to_anchor=(1.05, 1), loc="upper left")
- plt.tight_layout()
- plt.show()
- # %%
- # Step 1: Remove Group A
- samples_to_keep = pca_df[pca_df["Group"] != "A"]["Sample"]
- df_filtered = df_imputed.loc[:, samples_to_keep]
- # Step 2: Transpose and fix column types
- df_filtered_T = df_filtered.T.copy()
- df_filtered_T.columns = df_filtered_T.columns.astype(str)
- # Step 3: Standardize
- from sklearn.preprocessing import StandardScaler
- scaler = StandardScaler()
- df_scaled_filtered = scaler.fit_transform(df_filtered_T)
- # ✅ Step 4: Run PCA on filtered data
- from sklearn.decomposition import PCA
- pca = PCA(n_components=2)
- pca_data_filtered = pca.fit_transform(df_scaled_filtered) # 🔥 This was missing
- # ✅ Step 5: Run KMeans clustering on PCA output
- from sklearn.cluster import KMeans
- kmeans = KMeans(n_clusters=2, random_state=42)
- clusters = kmeans.fit_predict(pca_data_filtered)
- # Step 6: Build new PCA DataFrame
- import pandas as pd
- pca_filtered_df = pd.DataFrame(pca_data_filtered, columns=["PC1", "PC2"])
- pca_filtered_df["Sample"] = df_filtered.columns
- pca_filtered_df["Group"] = pca_filtered_df["Sample"].str[0]
- pca_filtered_df["Cluster"] = clusters # ✅ Now this works!
- # Step 7: Plot
- import matplotlib.pyplot as plt
- fig, ax = plt.subplots(figsize=(10, 8))
- for cluster in pca_filtered_df["Cluster"].unique():
- cluster_data = pca_filtered_df[pca_filtered_df["Cluster"] == cluster]
- ax.scatter(
- cluster_data["PC1"], cluster_data["PC2"],
- label=f"Cluster {cluster}", s=200, alpha=0.7
- )
- explained = pca.explained_variance_ratio_ * 100
- ax.set_title("PCA after Removing Group A", fontsize=16)
- ax.set_xlabel(f"PC1 ({explained[0]:.1f}%)", fontsize=14)
- ax.set_ylabel(f"PC2 ({explained[1]:.1f}%)", fontsize=14)
- ax.legend(title="KMeans Cluster")
- plt.tight_layout()
- plt.show()
- # %%
- print(f"Explained variance (PC1): {explained[0]:.2f}%")
- print(f"Explained variance (PC2): {explained[1]:.2f}%")
- # %%
- from scipy.stats import zscore
- # Step 1: Extract gene expression across all samples
- gene = "NES"
- gene_expr = df_imputed.loc[gene] # This is a Series: index = sample names, values = expression
- # Step 2: Convert to DataFrame and add Cluster info
- gene_df = pd.DataFrame({"Expression": gene_expr})
- gene_df["Sample"] = gene_df.index
- gene_df["Cluster"] = gene_df["Sample"].map(pca_df.set_index("Sample")["Cluster"])
- # Step 3: Compute Z-score of expression across all samples
- gene_df["Z_score"] = zscore(gene_df["Expression"])
- # Step 4: Compare distributions across clusters
- import seaborn as sns
- import matplotlib.pyplot as plt
- plt.figure(figsize=(8, 6))
- sns.boxplot(data=gene_df, x="Cluster", y="Z_score", palette="Set2")
- sns.stripplot(data=gene_df, x="Cluster", y="Z_score", color="black", alpha=0.6)
- plt.title(f"Z-score of {gene} Expression Across Clusters", fontsize=16)
- plt.xlabel("KMeans Cluster", fontsize=14)
- plt.ylabel("Z-score", fontsize=14)
- plt.tight_layout()
- plt.show()
- # Step 4: Export to Excel
- output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\NES_zscores.xlsx"
- #gene_df.to_excel(output_path, index=False)
- print("✅ Exported Z-scores to:", output_path)
- # %%
- from scipy.stats import zscore
- # Step 1: Define your list of genes (or use top ones)
- genes_of_interest = ["NES", "TUBB3", "GFAP", "PRPH", "SLC17A7", "SLC17A6", "RBFOX3", "MAP2", "SYN1","SOX2","CD44", "SYP", "MKI67", "ALDH1L1", "DLG4", "DCX"] # Replace with your list
- # Step 2: Transpose df_imputed to shape (samples x genes)
- df_expr = df_imputed.T.copy()
- df_expr["Sample"] = df_expr.index
- df_expr["Group"] = df_expr["Sample"].str[0] # Adjust if needed
- df_expr["Cluster"] = df_expr["Sample"].map(pca_df.set_index("Sample")["Cluster"])
- # Step 3: Melt into long format
- df_long = df_expr[genes_of_interest + ["Sample", "Group", "Cluster"]].melt(
- id_vars=["Sample", "Group", "Cluster"],
- var_name="Gene",
- value_name="Expression"
- )
- # Step 4: Compute Z-scores for each gene across all samples
- df_long["Z_score"] = df_long.groupby("Gene")["Expression"].transform(zscore)
- # Step 5: Export to Excel
- output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\multi_gene_zscores1.xlsx"
- #df_long.to_excel(output_path, index=False)
- print("✅ Exported multi-gene Z-scores to:", output_path)
- # %%
- import seaborn as sns
- import matplotlib.pyplot as plt
- # Use the DataFrame from your earlier export, e.g., `df_long`
- plt.figure(figsize=(14, 8))
- sns.boxplot(data=df_long, x="Group", y="Z_score", hue="Gene", palette="Set2")
- plt.title("Z-score of Multiple Genes by Group", fontsize=16)
- plt.xlabel("Group", fontsize=14)
- plt.ylabel("Z-score", fontsize=14)
- plt.legend(title="Gene", bbox_to_anchor=(1.05, 1), loc="upper left")
- plt.tight_layout()
- plt.show()
- # %%
- plt.figure(figsize=(14, 8))
- sns.violinplot(
- data=df_long,
- x="Gene",
- y="Z_score",
- hue="Cluster", # ✅ Switch hue to cluster instead
- palette="Set2",
- dodge=True # Separates the clusters side-by-side
- )
- plt.title("Violin Plot of Gene Z-scores by Cluster", fontsize=16)
- plt.xlabel("Gene")
- plt.ylabel("Z-score")
- plt.legend(title="Cluster", bbox_to_anchor=(1.05, 1), loc="upper left")
- plt.tight_layout()
- plt.show()
- # %%
- from scipy.stats import mannwhitneyu
- import seaborn as sns
- import matplotlib.pyplot as plt
- import os
- # Create output folder
- #output_dir = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\violin_plots"
- #os.makedirs(output_dir, exist_ok=True)
- # ✅ Define custom colors (do this outside the loop — one time)
- custom_palette = {0: "lightblue", 1: "lightgray"}
- # Loop through genes
- for gene in df_long["Gene"].unique():
- plt.figure(figsize=(6, 5))
- # Subset data
- df_gene = df_long[df_long["Gene"] == gene]
- group0 = df_gene[df_gene["Cluster"] == 0]["Z_score"]
- group1 = df_gene[df_gene["Cluster"] == 1]["Z_score"]
- # Mann–Whitney U
- u_stat, p_val = mannwhitneyu(group0, group1, alternative='two-sided')
- # Significance stars
- if p_val < 0.001:
- sig = "***"
- elif p_val < 0.01:
- sig = "**"
- elif p_val < 0.05:
- sig = "*"
- else:
- sig = "ns"
- # ✅ Plot using your custom colors
- sns.violinplot(
- data=df_gene,
- x="Cluster",
- y="Z_score",
- hue="Cluster",
- split=True,
- palette=custom_palette,
- dodge=False
- )
- plt.legend().remove()
- # Style
- plt.title(f"{gene} (Mann–Whitney p = {p_val:.3e}, {sig})", fontsize=16)
- plt.xlabel("Cluster", fontsize=16)
- plt.ylabel("Z-score", fontsize=16)
- plt.xticks(fontsize=16)
- plt.yticks(fontsize=16)
- # Thicken axes lines
- ax = plt.gca()
- for spine in ax.spines.values():
- spine.set_linewidth(1)
- # Major ticks
- ax.tick_params(axis='both', which='major', labelsize=16, length=6, width=1.5)
- # Add significance annotation
- y_max = df_gene["Z_score"].max()
- plt.ylim(top=y_max + 1)
- plt.plot([0.25, 0.75], [y_max + 0.3]*2, lw=1.5, c='black')
- plt.text(0.5, y_max + 0.5, sig, ha='center', va='center', fontsize=16)
- # Save as SVG
- output_path = os.path.join(output_dir, f"{gene}_violin_cluster.svg")
- #plt.savefig(output_path, format="svg", bbox_inches="tight")
- plt.tight_layout()
- plt.show()
- # %% [markdown]
- # ## For your volcano plot, you want to compare expression values (not Z-scores) between clusters.
- #
- # ## So:
- # ## ❌ Don't use df_scaled_df or df_scaled_T for volcano
- # ## ✅ Use the log-transformed, imputed, normalized data → df_imputed
- # %%
- # Already log-transformed, normalized, and imputed
- df_expression = df_imputed
- # %%
- df_imputed
- # %%
- sample_to_group = df_expression.columns.str[0] # assumes sample names start with group letter
- # %%
- group1_labels = ['A', 'B', 'C', 'D']
- group2_labels = ['E', 'F', 'G', 'H']
- sample_groups = df_expression.columns.to_series().str[0]
- group1_samples = sample_groups[sample_groups.isin(group1_labels)].index
- group2_samples = sample_groups[sample_groups.isin(group2_labels)].index
- # %%
- group1 = df_expression[group1_samples]
- group2 = df_expression[group2_samples]
- # %%
- from scipy.stats import ttest_ind
- import numpy as np
- import pandas as pd
- from statsmodels.stats.multitest import multipletests
- import matplotlib.pyplot as plt
- # Calculate means
- mean_group1 = group1.mean(axis=1)
- mean_group2 = group2.mean(axis=1)
- # Log2 fold change (already log-transformed)
- log2fc = mean_group2 - mean_group1
- # t-test
- pvals = ttest_ind(group2.T, group1.T, axis=0, equal_var=False).pvalue
- # FDR correction (Benjamini-Hochberg)
- rejected, pvals_corrected, _, _ = multipletests(pvals, alpha=0.05, method='fdr_bh')
- # Create volcano dataframe
- volcano_df = pd.DataFrame({
- "Gene": df_expression.index,
- "log2FC": log2fc,
- "pval": pvals,
- "pval_adj": pvals_corrected,
- "significant": rejected
- })
- volcano_df["neg_log10_pval"] = -np.log10(volcano_df["pval_adj"].clip(lower=1e-300))
- volcano_df = volcano_df.replace([np.inf, -np.inf], np.nan).dropna()
- # Volcano plot then I
- plt.figure(figsize=(10, 6))
- plt.scatter(
- volcano_df["log2FC"],
- volcano_df["neg_log10_pval"],
- c=volcano_df["significant"].map({True: "red", False: "gray"}),
- alpha=0.7
- )
- # Label all genes with abs(log2FC) > 1 and FDR < 0.05
- highlighted = volcano_df[(volcano_df["significant"]) & (volcano_df["log2FC"].abs() > 1)]
- for _, row in highlighted.iterrows():
- plt.text(row["log2FC"], row["neg_log10_pval"], row["Gene"],
- fontsize=5, ha='center', va='bottom')
- # Threshold lines
- plt.axhline(-np.log10(0.05), linestyle='--', color='gray')
- plt.axvline(1, linestyle='--', color='gray')
- plt.axvline(-1, linestyle='--', color='gray')
- plt.title("Volcano Plot with FDR Correction")
- plt.xlabel("Log₂ Fold Change")
- plt.ylabel("-Log₁₀ Adjusted p-value")
- plt.tight_layout()
- plt.show()
- # %% [markdown]
- # ## log2fc = mean_group2 - mean_group1
- # ##So:
- #
- # ##If log2FC > 0, the gene is higher in Group 2
- #
- # ##If log2FC < 0, it’s higher in Group 1
- # %%
- sig_genes_df = volcano_df[
- (volcano_df["pval_adj"] < 0.05) &
- (volcano_df["log2FC"].abs() > 1)
- ]
- # Sort by most significant
- sig_genes_df = sig_genes_df.sort_values("pval_adj")
- # View top 10
- print(sig_genes_df.head(10))
- # Export all to Excel
- sig_genes_df.to_excel("significant_genes_filtered.xlsx", index=False)
- # %%
- # Calculate means
- mean_group1 = group1.mean(axis=1)
- mean_group2 = group2.mean(axis=1)
- # Log2 fold change (already log-transformed)
- log2fc = mean_group2 - mean_group1
- # t-test
- pvals = ttest_ind(group2.T, group1.T, axis=0, equal_var=False).pvalue
- # FDR correction (Benjamini-Hochberg)
- rejected, pvals_corrected, _, _ = multipletests(pvals, alpha=0.05, method='fdr_bh')
- # Create volcano dataframe
- volcano_df = pd.DataFrame({
- "Gene": df_expression.index,
- "log2FC": log2fc,
- "pval": pvals,
- "pval_adj": pvals_corrected,
- "significant": rejected
- })
- volcano_df["neg_log10_pval"] = -np.log10(volcano_df["pval_adj"].clip(lower=1e-300))
- volcano_df = volcano_df.replace([np.inf, -np.inf], np.nan).dropna()
- # Volcano plot
- plt.figure(figsize=(10, 6))
- plt.scatter(
- volcano_df["log2FC"],
- volcano_df["neg_log10_pval"],
- c=volcano_df["significant"].map({True: "lightblue", False: "gray"}),
- alpha=0.7
- )
- # 🎯 Annotate specific genes of interest
- genes_to_annotate = ["SOX2", "NES", "MIK67", "POU3F", "NOTCH1", "SLC6A9", "GFAP", "GRM2", "GRM5", "RGS10", "FOXP1", "CD44", "ALDH1L1", "SYT7", "GLUL", "SYT11", "SCGN", "SLC6A1", "SLC6A9" "NOTCH1", "DCX", "SDC1", "EGFR", "PHPR"] # ← customize this
- # 1. Not significant
- df_gray = volcano_df[~volcano_df["significant"]]
- # 2. Significant but not in your gene list
- df_lightblue = volcano_df[
- (volcano_df["significant"]) & (~volcano_df["Gene"].isin(genes_to_annotate))
- ]
- # 3. Highlighted genes (e.g., NES, TP53, etc.)
- df_pink = volcano_df[volcano_df["Gene"].isin(genes_to_annotate)]
- plt.figure(figsize=(10, 6))
- ax = plt.gca()
- # Plot gray first
- ax.scatter(df_gray["log2FC"], df_gray["neg_log10_pval"],
- color="gray", alpha=0.6, s=30, label="Not significant")
- # Plot lightblue next
- ax.scatter(df_lightblue["log2FC"], df_lightblue["neg_log10_pval"],
- color="lightblue", alpha=0.7, s=30, label="Significant")
- # Plot pink last (so it’s on top)
- ax.scatter(df_pink["log2FC"], df_pink["neg_log10_pval"],
- color="bisque", edgecolor="darkorange", linewidth=1.0,
- s=60, label="Highlighted Genes")
- for _, row in df_pink.iterrows():
- ax.text(row["log2FC"], row["neg_log10_pval"], row["Gene"],
- fontsize=10, ha='center', va='bottom', fontweight='bold', color='black')
- plt.axhline(-np.log10(0.05), linestyle='--', color='gray')
- plt.axvline(1, linestyle='--', color='gray')
- plt.axvline(-1, linestyle='--', color='gray')
- plt.xlabel("Log₂ Fold Change", fontsize=16)
- plt.ylabel("−Log₁₀ Adjusted p-value", fontsize=16)
- plt.title("Volcano Plot with Layered Highlighting", fontsize=16)
- #plt.legend(frameon=True)
- plt.tight_layout()
- #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\Volcano plot.svg", format="svg")
- plt.show()
- # %%
- # 1. Filter volcano_df for significant DE genes
- sig_genes_df = volcano_df[
- (volcano_df["pval_adj"] < 0.05) &
- (volcano_df["log2FC"].abs() > 1)
- ].copy()
- # 2. Sort by adjusted p-value
- sig_genes_df = sig_genes_df.sort_values("pval_adj")
- # 3. (Optional) View top 10
- print("Top significant genes:\n", sig_genes_df.head(10))
- # 4. Export to Excel
- output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\sig_genes.xlsx"
- #sig_genes_df.to_excel(output_path, index=False)
- print(f"Exported to: {output_path}")
- # %%
- #!pip install gprofiler-official
- # %%
- from gprofiler import GProfiler
- # Get your gene list
- gene_list = sig_genes_df["Gene"].tolist()
- # Create gProfiler object
- gp = GProfiler(return_dataframe=True)
- # Run GO/pathway enrichment
- results = gp.profile(
- organism="hsapiens",
- query=gene_list,
- user_threshold=0.05,
- sources=["GO:BP", "GO:MF", "GO:CC", "KEGG", "REAC", "WP"]
- )
- # Step 4: Check results
- print(f"Submitted {len(gene_list)} genes.")
- if results.empty:
- print("⚠️ No significant enrichment found. Try relaxing thresholds or using more genes.")
- else:
- print("✅ Enrichment completed. Columns available:", results.columns.tolist())
- # Only show if 'term_name' exists
- if "term_name" in results.columns:
- print(results[["term_name", "p_value", "source"]].head(10))
- else:
- print("⚠️ 'term_name' column not found in results. Full results:")
- print(results.head())
- # Save results
- output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\gprofiler_enrichment_results.xlsx"
- #results.to_excel(output_path, index=False)
- print(f"✅ Saved results to {output_path}")
- # %%
- print("Columns in results:", results.columns.tolist())
- print(results.head())
- # %%
- print(f"Number of enriched terms found: {results.shape[0]}")
- # %%
- import seaborn as sns
- import matplotlib.pyplot as plt
- # Make sure results aren't empty
- if not results.empty and "name" in results.columns:
- # Step 1: Pick top N terms (you can change 10)
- top_terms = results.sort_values("p_value").head(10)
- # Step 2: Barplot of top pathways
- plt.figure(figsize=(8, 6))
- sns.barplot(
- y="name",
- x="-log10(p_value)",
- data=top_terms.assign(**{"-log10(p_value)": -np.log10(top_terms["p_value"])}),
- palette="viridis"
- )
- plt.title("Top 10 Enriched Pathways (g:Profiler)", fontsize=16)
- plt.xlabel("-Log₁₀ p-value", fontsize=14)
- plt.ylabel("Pathway", fontsize=14)
- plt.tight_layout()
- plt.show()
- else:
- print("⚠️ No results to plot.")
- # %%
- # Find the row for the pathway you care about
- row = results[results["name"] == "cell junction"]
- if not row.empty:
- print("Matched Genes:", row["intersection_size"].values[0])
- else:
- print("⚠️ No pathway found with that exact name!")
- # Show the genes (usually in 'intersection')
- print(row["intersection_size"].values[0])
- # %%
- import seaborn as sns
- import matplotlib.pyplot as plt
- import numpy as np
- # Confirm 'results' is populated and has expected columns
- if not results.empty and "name" in results.columns and "p_value" in results.columns:
- # Step 1: Compute -log10(p-value)
- results["log_p"] = -np.log10(results["p_value"])
- # Step 2: Get top 10 terms by significance
- top_terms = results.nsmallest(10, "p_value")
- # Step 3: Bubble plot
- plt.figure(figsize=(10, 8))
- sns.scatterplot(
- data=top_terms,
- x="intersection_size", # Number of DEGs in each term
- y="name", # GO term or pathway name
- size="intersection_size",
- hue="log_p",
- sizes=(100, 1000),
- palette="viridis",
- legend="brief"
- )
- plt.xlabel("Gene Count in Term", fontsize=12)
- plt.ylabel("Pathway / GO Term", fontsize=12)
- plt.title("Top Enriched Terms (Bubble Plot)", fontsize=14)
- plt.tight_layout()
- plt.show()
- else:
- print("⚠️ No valid enrichment results found to plot.")
- # %%
- import matplotlib.pyplot as plt
- import seaborn as sns
- import numpy as np
- # List of GO terms or pathway names you're interested in (case-insensitive match)
- terms_of_interest = [
- "cell differentiation",
- "astrocyte differentiation",
- "nervous system development",
- "neurogenesis",
- "generation of neurons",
- "neuron differentiation",
- "cortical cytoskeleton organization",
- "synapse",
- "axon development",
- "gliogenesis",
- "Neuroinflammation and glutamatergic signaling",
- ]
- # --- 1. Prepare p-value for color scale ---
- filtered_terms["p_value_capped"] = filtered_terms["p_value"].clip(lower=1e-10) # avoid log(0)
- filtered_terms["p_label"] = filtered_terms["p_value"].apply(lambda x: f"{x:.1e}")
- # Sort to control y-axis order (optional)
- filtered_terms = filtered_terms.sort_values("intersection_size", ascending=True).reset_index(drop=True)
- # --- Calculate plot boundaries safely ---
- max_x = filtered_terms["intersection_size"].max()
- x_offset = max_x * 0.40 # shift p-values to the right by ~10%
- x_buffer = max_x * 0.25 # add buffer to x-axis so text isn't cut off
- # --- Set up plot ---
- fig, ax = plt.subplots(figsize=(12, 8))
- bubble = sns.scatterplot(
- data=filtered_terms,
- x="intersection_size",
- y="name",
- size="intersection_size",
- hue="p_value",
- palette="Blues_r",
- sizes=(200, 1200),
- alpha=0.8,
- edgecolor="gray",
- legend="brief",
- ax=ax
- )
- # --- Add p-value text BELOW bubble ---
- for i, row in filtered_terms.iterrows():
- ax.text(
- row["intersection_size"],
- i - 0.5, # further down
- row["p_label"],
- ha="left",
- va="top",
- fontsize=12,
- color="black"
- )
- # --- Adjust font sizes ---
- ax.set_yticks(range(len(filtered_terms)))
- ax.set_yticklabels(filtered_terms["name"], fontsize=14)
- ax.tick_params(axis='x', labelsize=12)
- ax.set_xlabel("Gene Count in Term", fontsize=14)
- ax.set_ylabel("GO Term", fontsize=14)
- ax.set_title("Selected GO Terms (Enrichment Bubble Plot)", fontsize=16)
- # --- Prevent clipping of bubbles or labels ---
- ax.set_xlim(0, max_x + x_buffer)
- ax.set_ylim(-1, len(filtered_terms) - 0.5)
- # --- Fix legend position outside plot ---
- sns.move_legend(bubble, "upper left", bbox_to_anchor=(1.02, 1), frameon=True)
- # --- Improve layout ---
- plt.grid(axis="x", linestyle="--", alpha=0.3)
- plt.tight_layout(rect=[0, 0, 1, 1]) # leave space on right for legend
- #plt.savefig(r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\Enrichment Bubble Plot.svg", format="svg", bbox_inches="tight")
- plt.show()
- # %%
- # Example: Plot a single protein across conditions
- proteins_of_interest = ["NES", "MAP2", "GFAP"] # Replace with your protein
- # Make sure your expression data is in long format
- # Melt data for seaborn plotting
- expr_long = df_imputed.loc[proteins_of_interest].T.reset_index().melt(id_vars="index", var_name="Protein", value_name="Expression")
- expr_long.rename(columns={"index": "Sample"}, inplace=True)
- # Add group info from sample names if encoded
- expr_long["Group"] = expr_long["Sample"].str[0] # Adjust if needed
- # Plot
- plt.figure(figsize=(12, 6))
- sns.boxplot(data=expr_long, x="Protein", y="Expression", hue="Group")
- sns.stripplot(data=expr_long, x="Protein", y="Expression", hue="Group", dodge=True, color='black', size=4, alpha=0.5)
- plt.title("Expression of Selected Proteins by Group")
- plt.xticks(rotation=45)
- plt.tight_layout()
- plt.show()
- # %% [markdown]
- # # df_expression = df_imputed # ⬅️ This was your imputed expression matrix df_imputed = df_expression # If df_expression is already defined
- # %%
- df_imputed = df_expression # If df_expression is already defined
- # %%
- df
- # %%
- # Assuming your DataFrame is named 'df'
- column_sums = df.sum(axis=0) # Calculate the sum of each column
- # Divide each element by the corresponding column sum
- df_normalized = df.div(column_sums, axis=1)
- df_normalized
- # %%
- # Handle missing values (e.g., impute with mean or median), imputation may contain very small numbers
- # Create a KNN imputer object
- imputer = KNNImputer(n_neighbors=5) # Adjust n_neighbors as needed
- # Impute missing values using KNN
- df_imputed = imputer.fit_transform(df_normalized)
- # Create a new DataFrame with imputed values
- df_imputed = pd.DataFrame(df_imputed, columns=df.columns, index= df.index)
- df_imputed*10e6
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import seaborn as sns
- from scipy.stats import mannwhitneyu
- import os
- import math
- # --- Scale expression data ---
- df_scaled = df_imputed * 1e7
- # --- Assign groups from sample names ---
- group1_labels = ['A', 'B', 'C', 'D']
- group2_labels = ['E', 'F', 'G', 'H']
- sample_to_group = df_scaled.columns.str[0]
- group_map = sample_to_group.map(lambda x: 'Group 1' if x in group1_labels else ('Group 2' if x in group2_labels else 'Other'))
- # --- Convert to long-form dataframe ---
- df_long = df_scaled.T.copy()
- df_long["Group"] = group_map.values
- df_long = df_long.reset_index().melt(id_vars=["index", "Group"], var_name="Gene", value_name="Expression")
- df_long.rename(columns={"index": "Sample"}, inplace=True)
- # --- Genes of interest ---
- genes_of_interest = ["STAT3", "ZEB2", "GFAP", "PLPP3", "SOX9", "NR3C1"] # add more as needed
- n_genes = len(genes_of_interest)
- # --- Grid dimensions (2 columns) ---
- n_cols = 2
- n_rows = math.ceil(n_genes / n_cols)
- # --- Set up the figure ---
- fig, axes = plt.subplots(n_rows, n_cols, figsize=(14, 5 * n_rows), squeeze=False)
- # --- Plot each gene in grid ---
- for idx, gene in enumerate(genes_of_interest):
- row = idx // n_cols
- col = idx % n_cols
- ax = axes[row][col]
- df_gene = df_long[df_long["Gene"] == gene]
- group1 = df_gene[df_gene["Group"] == "Group 1"]["Expression"]
- group2 = df_gene[df_gene["Group"] == "Group 2"]["Expression"]
- # Mann–Whitney U test
- u_stat, p_val = mannwhitneyu(group1, group2, alternative="two-sided")
- sig = "***" if p_val < 0.001 else "**" if p_val < 0.01 else "*" if p_val < 0.05 else "ns"
- # Plot
- sns.violinplot(data=df_gene, x="Group", y="Expression", palette=["lightblue", "lightgray"], ax=ax, inner=None)
- sns.stripplot(data=df_gene, x="Group", y="Expression", color="black", size=4, jitter=True, alpha=0.6, ax=ax)
- # Significance annotation
- y_max = df_gene["Expression"].max()
- ax.set_ylim(top=y_max + 0.25 * y_max)
- ax.plot([0, 1], [y_max + 0.05 * y_max] * 2, lw=1.5, c="black")
- ax.text(0.5, y_max + 0.1 * y_max, sig, ha="center", va="center", fontsize=14)
- # Style
- ax.set_title(f"{gene} (p = {p_val:.2e}, {sig})", fontsize=13)
- ax.set_xlabel("")
- ax.set_ylabel("Expression", fontsize=12)
- ax.tick_params(axis="both", labelsize=11)
- # --- Hide unused subplots if total isn't a full grid ---
- for i in range(n_genes, n_rows * n_cols):
- fig.delaxes(axes[i // n_cols][i % n_cols])
- # --- Final layout and export ---
- plt.tight_layout()
- output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\violin_plots\multi_gene_violin_grid.svg"
- #plt.savefig(output_path, format="svg", bbox_inches="tight")
- plt.show()
- # %%
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- import seaborn as sns
- from scipy.stats import mannwhitneyu
- # --- List of genes to analyze ---
- genes_of_interest = ["STAT3", "ZEB2", "GFAP", "SOX9", "NR3C1", "NOTCH1", "QKI", "SYN1", "DLG4", "SLC17A7", "SLC17A6", "CCNE1", "CCNA2", "DCX", "EGFR", "SDC1", "SOX2", "PRPH"] # customize as needed
- # --- Setup for results ---
- results = []
- for gene in genes_of_interest:
- group1 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['A', 'B', 'C', 'D'])]
- group2 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['E', 'F', 'G', 'H'])]
- # Compute log2 fold change
- fc = group1.mean() / group2.mean()
- log2fc = np.log2(fc)
- # Mann–Whitney U test
- u_stat, p_val = mannwhitneyu(group1, group2, alternative='two-sided')
- # Significance label
- if p_val < 0.001:
- sig = '***'
- elif p_val < 0.01:
- sig = '**'
- elif p_val < 0.05:
- sig = '*'
- else:
- sig = 'ns'
- # Collect
- results.append({"Gene": gene, "log2FC": log2fc, "pval": p_val, "sig": sig})
- # --- Convert to DataFrame ---
- fc_df = pd.DataFrame(results)
- # --- Plot ---
- plt.figure(figsize=(10, 6))
- sns.barplot(data=fc_df, x="Gene", y="log2FC", color="lightblue", edgecolor="gray")
- # Add significance stars above bars
- for i, row in fc_df.iterrows():
- y = row["log2FC"]
- offset = 0.05 * np.max(np.abs(fc_df["log2FC"]))
- plt.text(i, y + np.sign(y) * offset, row["sig"], ha='center', va='bottom' if y >= 0 else 'top', fontsize=14)
- # Style
- plt.axhline(0, color="black", lw=1)
- plt.ylabel("log₂ Fold Change (Group1 / Group2)", fontsize=13)
- plt.xlabel("")
- plt.title("Fold Change Between Groups with Expression Significance", fontsize=14)
- plt.xticks(fontsize=12)
- plt.yticks(fontsize=12)
- plt.tight_layout()
- # Expand y-limits so stars fit
- y_max = fc_df["log2FC"].max()
- y_min = fc_df["log2FC"].min()
- y_range = y_max - y_min
- plt.ylim(y_min - 0.2 * y_range, y_max + 0.2 * y_range)
- output_path = r"C:\Users\MunozestraJ\OneDrive - Cedars-Sinai Health System\Proteomics\Dox experiment\foldchange astrocytes differentiation.svg"
- #plt.savefig(output_path, format="svg", bbox_inches="tight")
- plt.show()
- # %%
- # Display sorted fold change table
- print(fc_df.sort_values("log2FC", ascending=False))
- # %%
- import numpy as np
- import pandas as pd
- import matplotlib.pyplot as plt
- import seaborn as sns
- from scipy.stats import mannwhitneyu
- # --- List of genes to analyze ---
- #genes_of_interest = ["SLC6A9", "GRM2", "PSAT1", "SHMT1", "GLUL", "GRM5", "SLC1A3", "GRIK1"] # customize as needed
- genes_of_interest = ["STAT3", "ZEB2", "GFAP", "SOX9", "NR3C1", "NOTCH1", "QKI"]
- # --- Setup for results ---
- results = []
- for gene in genes_of_interest:
- group1 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['A', 'B', 'C', 'D'])]
- group2 = df_scaled.loc[gene, df_scaled.columns.str[0].isin(['E', 'F', 'G', 'H'])]
- # Compute log2 fold change
- fc = group2.mean() / group1.mean()
- log2fc = np.log2(fc)
- # Mann–Whitney U test
- u_stat, p_val = mannwhitneyu(group1, group2, alternative='two-sided')
- # Significance label
- if p_val < 0.001:
- sig = '***'
- elif p_val < 0.01:
- sig = '**'
- elif p_val < 0.05:
- sig = '*'
- else:
- sig = 'ns'
- # Collect
- results.append({"Gene": gene, "log2FC": log2fc, "pval": p_val, "sig": sig})
- # --- Convert to DataFrame ---
- fc_df = pd.DataFrame(results)
- # --- Prepare plot ---
- plt.figure(figsize=(10, 6))
- # Sort genes by fold change for visual clarity
- fc_df_sorted = fc_df.sort_values("log2FC", ascending=False).reset_index(drop=True)
- # Plot dots (log2FC on y-axis)
- sns.stripplot(
- data=fc_df_sorted,
- x="log2FC",
- y="Gene",
- size=10,
- color="lightblue",
- edgecolor="orange",
- )
- # Add vertical line at 0 (no change)
- plt.axvline(0, color="black", lw=1)
- # Add significance stars next to dots
- x_range = fc_df_sorted["log2FC"].max() - fc_df_sorted["log2FC"].min()
- x_offset = 0.03 * x_range # offset stars slightly to the right
- for i, row in fc_df_sorted.iterrows():
- x = row["log2FC"]
- y = i
- plt.text(
- x + x_offset,
- y,
- row["sig"],
- va="center",
- ha="left",
- fontsize=13,
- color="black"
- )
- # Style
- plt.xlabel("log₂ Fold Change (Group1 / Group2)", fontsize=13)
- plt.ylabel("")
- plt.title("Fold Change by Gene with Significance", fontsize=14)
- plt.xticks(fontsize=12)
- plt.yticks(fontsize=12)
- plt.tight_layout()
- plt.show()
- # %%
- import pandas as pd
- import seaborn as sns
- import matplotlib.pyplot as plt
- from scipy.stats import ttest_ind
- # --- Protein list to plot ---
- proteins_of_interest = ["STAT3", "ZEB2", "GFAP", "PLPP3", "SOX9", "NR3C1"] # Update as needed
- # --- Melt expression data ---
- expr_long = df_imputed.loc[proteins_of_interest].T.reset_index().melt(
- id_vars="index", var_name="Protein", value_name="Expression"
- )
- expr_long.rename(columns={"index": "Sample"}, inplace=True)
- # --- Assign groups using first letter of sample name ---
- expr_long["Group"] = expr_long["Sample"].str[0]
- # --- Plot: One violin plot per protein ---
- plt.figure(figsize=(12, 6 * len(proteins_of_interest)))
- for i, protein in enumerate(proteins_of_interest):
- plt.subplot(len(proteins_of_interest), 1, i + 1)
- # Subset data for this protein
- data = expr_long[expr_long["Protein"] == protein]
- # Group split
- group_values = [g["Expression"].values for _, g in data.groupby("Group")]
- if len(group_values) == 2:
- stat, pval = ttest_ind(*group_values)
- else:
- pval = float('nan')
- # Determine significance label
- if pval < 0.001:
- sig_label = "***"
- elif pval < 0.01:
- sig_label = "**"
- elif pval < 0.05:
- sig_label = "*"
- else:
- sig_label = "ns"
- # Plot violin and strip
- sns.violinplot(data=data, x="Group", y="Expression", inner=None, palette="pastel")
- sns.stripplot(data=data, x="Group", y="Expression", color="black", size=4, alpha=0.6)
- # Annotate significance
- y_max = data["Expression"].max()
- y_min = data["Expression"].min()
- y_range = y_max - y_min
- plt.text(0.5, y_max + 0.05 * y_range, sig_label, ha="center", fontsize=14)
- # Titles and labels
- plt.title(f"{protein} (p = {pval:.2e})")
- plt.xlabel("Condition")
- plt.ylabel("Expression")
- plt.tight_layout()
- plt.show()
- # %%
- # %%
- # %%
- # %%
PCA and Volcano plot.ipynb at commit 9be7e24, no license · at the source
Overview
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repository
Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.
xomicsdatascience/generation_of_iNeurons
9be7e24aeb13dcdd0470bc641b06cf1d171ab335, 30 January 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
11 files
- python/
2025-ipscs-scp-github.ip , Jupyter, 296 lines, 1 matchynb - python/
Baseline linear model .ipynb , Jupyter, 383 lines, 1 match - python/
Cluster heatmap and centroid distances.ipynb , Jupyter, 439 lines, 1 match - python/
Neuronal and non neuronal.ipynb , Jupyter, 238 lines - python/
PCA and Volcano plot.ipynb , Jupyter, 1,620 lines, 4 matches - python/
PCA_clustering_heatmap.i , Jupyter, 207 linespynb - python/
clustering_correlation.i , Jupyter, 376 linespynb - python/
dataset_quality_geneplot , Jupyter, 547 lines.ipynb - python/
dataset_quality_geneplot , Jupyter, 632 lines2.ipynb - python/
linear model analysis visualization.ipynb , Jupyter, 388 lines - README.md, Text, 15 lines
Code availability statement
The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: xomicsdatascience/
generation_of_iNeurons
Read it in the paper: doi.org/10.1016/j.mcpro.2026.101604.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 10 scripts, each with its path and the digest of its content;
- 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability statement
The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- no repository, dataset or request procedure was recognized in it
Read it in the paper: doi.org/10.1016/j.mcpro.2026.101604.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Authors: added Jesse G Meyer (0000-0003-2753-3926); removed Jesse G Meyer
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 10 keywords, 13 MeSH terms, 4 funders, 37 references.
Cite
This paper
Muñoz-Estrada, J., Mostafania, A., Halwatura, L., Haghani, A., Jiang, Y., & Meyer, J. G. (2026). Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures. Molecular & cellular proteomics : MCP, 25(7), 101604. https://
BibTeX
@article{munozestrada202
author = {Muñoz-Estrada, Jesús and Mostafania, Andrew and Halwatura, Lahiruni and Haghani, Ali and Jiang, Yuming and Meyer, Jesse G},
title = {{Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures}},
journal = {Molecular \& cellular proteomics : MCP},
year = {2026},
month = jun,
volume = {25},
number = {7},
pages = {101604},
publisher = {American Society for Biochemistry and Molecular Biology},
issn = {1535-9476},
doi = {10.1016/
url = {https://
pmid = {42309337},
pmcid = {PMC13382583}
}
RIS
TY - JOUR
AU - Muñoz-Estrada, Jesús
AU - Mostafania, Andrew
AU - Halwatura, Lahiruni
AU - Haghani, Ali
AU - Jiang, Yuming
AU - Meyer, Jesse G
TI - Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures
T2 - Molecular & cellular proteomics : MCP
J2 - Mol Cell Proteomics
PY - 2026
DA - 2026/
VL - 25
IS - 7
SP - 101604
SN - 1535-9476
PB - American Society for Biochemistry and Molecular Biology
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "Optimizing NGN2 Dosage Enhances the Neuronal Enrichment of iPSC-Derived Neuronal Cultures",
"container-title": "Molecular & cellular proteomics : MCP",
"author": [
{
"family": "Muñoz-Estrada",
"given": "Jesús"
},
{
"family": "Mostafania",
"given": "Andrew"
},
{
"family": "Halwatura",
"given": "Lahiruni"
},
{
"family": "Haghani",
"given": "Ali"
},
{
"family": "Jiang",
"given": "Yuming"
},
{
"family": "Meyer",
"given": "Jesse G"
}
],
"container-title-short":
"volume": "25",
"issue": "7",
"page": "101604",
"DOI": "10.1016/
"PMID": "42309337",
"PMCID": "PMC13382583",
"ISSN": "1535-9476",
"publisher": "American Society for Biochemistry and Molecular Biology",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
17
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1093/nar/gkag368 [code]
- Single-cell trajectory inference for detecting transient events in biological processes.Journal: Nucleic acids researchIn common: anndata, Scanpy, scikit-learn, 4 other tools, genetics / omics, cellular / molecular, 3 references, author Jesse G Meyer
- [2] doi:10.1038/s41467-026-71803-3 [code]
- Charting the transition from in vitro gliogenesis to the in vivo maturation of human glial progenitor cells transplanted into the hypomyelinated mouse brain.Journal: Nature communicationsIn common: anndata, Scanpy, Plotly, 6 other tools, genetics / omics, cellular / molecular, 2 references
- [3] doi:10.1038/s41514-026-00391-9 [code]
- Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.Journal: npj agingIn common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics, cellular / molecular, 1 reference
- [4] doi:10.1038/s41593-026-02267-3 [code]
- Spatial proteomic analysis in human Alzheimer's disease brains enables identification of microenvironment-depende
nt microglial cell states. Journal: Nature neuroscienceIn common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics, cellular / molecular - [5] doi:10.1126/sciadv.aeg3223 [code]
- The extreme diversity of retinal amacrine cells has deep evolutionary roots.Journal: Science advancesIn common: anndata, Scanpy, Plotly, 6 other tools, genetics / omics, cellular / molecular, 1 reference
- [6] doi:10.1371/journal.pcbi.1014346 [code]
- StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion.Journal: PLoS computational biologyIn common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics
- [7] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics
- [8] doi:10.1016/j.celrep.2026.117270 [code]
- Blocking apoptosis promotes survival and alters developmental dynamics of human retinal ganglion cells in retinal organoids.Journal: Cell reportsIn common: anndata, Scanpy, Plotly, 7 other tools, genetics / omics
- [9] doi:10.3390/ijms27167275 [code]
- Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling.Journal: International journal of molecular sciencesIn common: anndata, Scanpy, statsmodels, 6 other tools, genetics / omics, cellular / molecular, 1 reference
- [10] doi:10.1038/s41467-026-76054-w [code]
- Distinct differentiation trajectories leave lasting impacts on gene regulation and function of V2a neurons.Journal: Nature communicationsIn common: Scanpy, Plotly, pandas, 2 other tools, cellular / molecular, 4 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 10 scripts, and 7 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:a1bb16104b612866…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
