Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease.
The 7 matches
- [1] § Results › Differential expression analysis ↔ src/code_for_lioness_networks_analysis/Sankey_Plot/plots(box_plot_sankey_plots_network_plots).ipynb, lines 1261–1262 · score 0.90 · TNF alpha, sodium symporter activity, hormone activity, extracellular matrix, amino acid, microvillus
- [2] § Results › Network metrics differences in AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/nci_netwoks_features_calculation.ipynb, lines 35–83 · score 0.71 · network properties, network metrics, connected component, largest, regulatory networks, modularity
- [3] § Results › Network metrics differences in AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/ad_netwoks_features_calculation.ipynb, lines 58–107 · score 0.71 · network properties, network metrics, connected component, largest, regulatory networks, modularity
- [4] § Methods › Pathway and cell type-specific enrichment analysis ↔ src/code_for_lioness_networks_analysis/Sankey_Plot/plots(box_plot_sankey_plots_network_plots).ipynb, lines 389–436 · score 0.67 · GO Cellular Component, GO Molecular Function, KEGG, Hallmark, Human, genes
- [5] § Methods › Network analysis of AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/nci_netwoks_features_calculation.ipynb, lines 35–83 · score 0.62 · network metrics, properties, largest, regulatory networks, modularity, nestedness
- [6] § Methods › Network analysis of AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/ad_netwoks_features_calculation.ipynb, lines 58–107 · score 0.62 · network metrics, properties, largest, regulatory networks, modularity, nestedness
- [7] § Methods › Data preprocessing ↔ src/code_for_data_preprocessing/combat_seq_batch_removal.ipynb, lines 1–15 · score 0.53 · Combat Seq, batch, preprocessing
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 1,329 lines · 36 KB · no license · 2 matches
- # %%
- import pandas as pd
- import os
- import networkx as nx
- from tqdm import tqdm
- # %%
- preprocessed_ad_single_sample_network_dir_path = '../../../data_preprocessed/grn_output/ad_nci_seperate_network/lioness/ad_preprocessed/'
- preprocessed_nci_single_sample_network_dir_path = '../../../data_preprocessed/grn_output/ad_nci_seperate_network/lioness/nci_preprocessed/'
- # %%
- AD_data = pd.read_csv(preprocessed_ad_single_sample_network_dir_path +'lioness.427_120507.58.csv').rename(columns={'2':'lioness.427_120507.58.csv'.split('.')[1]})
- for i in tqdm(os.listdir(preprocessed_ad_single_sample_network_dir_path)):
- if i.split('.')[1] in AD_data.columns:
- continue
- if 'csv' in i:
- df = pd.read_csv(preprocessed_ad_single_sample_network_dir_path + i).rename(columns={'2':i.split('.')[1]})
- AD_data = pd.merge(AD_data,df, on=['0','1'])
- # %%
- NCI_data = pd.read_csv(preprocessed_nci_single_sample_network_dir_path +'lioness.713_120531.47.csv').rename(columns={'2':'lioness.713_120531.47.csv'.split('.')[1]})
- for i in tqdm(os.listdir(preprocessed_nci_single_sample_network_dir_path)):
- if i.split('.')[1] in NCI_data.columns:
- continue
- if 'csv' in i:
- df = pd.read_csv(preprocessed_nci_single_sample_network_dir_path + i).rename(columns={'2':i.split('.')[1]})
- NCI_data = pd.merge(NCI_data,df, on=['0','1'])
- # %%
- AD_data.index = AD_data['0'] + '_' + AD_data['1']
- AD_data.drop(columns=['0','1'],inplace = True)
- AD_data
- # %%
- NCI_data.index = NCI_data['0'] + '_' + NCI_data['1']
- NCI_data.drop(columns=['0','1'],inplace = True)
- NCI_data
- # %%
- common_edge = set(NCI_data.index).intersection(set(AD_data.index))
- AD_data = AD_data.loc[list(common_edge),:]
- NCI_data = NCI_data.loc[list(common_edge),:]
- # %%
- AD_data = AD_data.T
- AD_data
- # %%
- NCI_data = NCI_data.T
- NCI_data
- # %%
- combined = pd.concat([AD_data,NCI_data])
- # %%
- Label = ['AD' for i in range(len(AD_data))] + ['NCI' for i in range(len(NCI_data))]
- Label
- # %%
- import pandas as pd
- from sklearn.decomposition import PCA
- import matplotlib.pyplot as plt
- import seaborn as sns
- # Generate sample data (replace this with your own DataFrame)
- combined['Label'] = Label
- # Assuming your DataFrame contains only numeric values
- # If not, you might need to preprocess your data accordingly
- # Perform PCA
- pca = PCA(n_components=2)
- principal_components = pca.fit_transform(combined.drop(columns = 'Label'))
- # Create a new DataFrame with the principal components
- pc_df = pd.DataFrame(data=principal_components, columns=['PC1', 'PC2'])
- # Concatenate the principal components DataFrame with the original DataFrame
- final_df = pd.concat([pc_df], axis=1)
- final_df['Label'] = Label
- final_df.index = combined.index
- # Plot the PCA results
- plt.figure(figsize=(8, 8))
- sns.scatterplot(x='PC1', y='PC2', data=final_df, hue = Label)
- plt.title('PCA Plot')
- plt.xlabel('Principal Component 1 (PC1)')
- plt.ylabel('Principal Component 2 (PC2)')
- plt.show()
- # %%
- remove = final_df[final_df.PC1> 80].index.tolist()
- remove
- # %%
- outlier_corrected_combined = combined[~combined.index.isin(remove)]
- # %%
- import pandas as pd
- from sklearn.decomposition import PCA
- import matplotlib.pyplot as plt
- import seaborn as sns
- # Generate sample data (replace this with your own DataFrame)
- # Assuming your DataFrame contains only numeric values
- # If not, you might need to preprocess your data accordingly
- # Perform PCA
- pca = PCA(n_components=2)
- principal_components = pca.fit_transform(outlier_corrected_combined .drop(columns = 'Label'))
- # Create a new DataFrame with the principal components
- pc_df = pd.DataFrame(data=principal_components, columns=['PC1', 'PC2'])
- # Concatenate the principal components DataFrame with the original DataFrame
- final_df = pd.concat([pc_df], axis=1)
- final_df['Label'] = outlier_corrected_combined.Label.tolist()
- final_df.index = outlier_corrected_combined.index
- Label = outlier_corrected_combined.Label.tolist()
- # Plot the PCA results
- plt.figure(figsize=(8, 8))
- sns.scatterplot(x='PC1', y='PC2', data=final_df, hue = Label )
- plt.title('PCA Plot')
- plt.xlabel('Principal Component 1 (PC1)')
- plt.ylabel('Principal Component 2 (PC2)')
- plt.show()
- # %%
- l = [
- 'ENSG00000166823_ENSG00000178199',
- 'ENSG00000157557_ENSG00000121903',
- 'ENSG00000028277_ENSG00000039139',
- 'ENSG00000081059_ENSG00000255561',
- 'ENSG00000167981_ENSG00000143222',
- 'ENSG00000198482_ENSG00000156219',
- 'ENSG00000188786_ENSG00000083799',
- 'ENSG00000168874_ENSG00000139197',
- 'ENSG00000139800_ENSG00000150175',
- 'ENSG00000079432_ENSG00000214425',
- 'ENSG00000162992_ENSG00000139197',
- 'ENSG00000256294_ENSG00000272576',
- 'ENSG00000169635_ENSG00000234841',
- 'ENSG00000125850_ENSG00000198783',
- 'ENSG00000078900_ENSG00000062096',
- 'ENSG00000120837_ENSG00000123179',
- 'ENSG00000079432_ENSG00000188234',
- 'ENSG00000188786_ENSG00000162384',
- 'ENSG00000251369_ENSG00000265018',
- 'ENSG00000234284_ENSG00000149089',
- 'ENSG00000198429_ENSG00000107262',
- 'ENSG00000188785_ENSG00000122641']
- # %%
- import pickle
- # Path to your pickle file
- pickle_file_path = '../../../data_preprocessed/pickle_file/ensemble_to_gene_name.pickle'
- # Load the pickle file
- with open(pickle_file_path, 'rb') as file:
- ensemble_to_gene_name = pickle.load(file)
- # %%
- def return_gene_name(gene):
- try:
- if ensemble_to_gene_name[gene] == 'Gene name not found':
- return gene
- return ensemble_to_gene_name[gene]
- except:
- return gene
- # %%
- n = []
- for i in l:
- m = []
- for j in i.split('_'):
- m.append(return_gene_name(j))
- n.append(m[0] + '-' + m[1])
- # %%
- n
- # %%
- l.append('Label')
- l
- # %%
- important_gene_dataframe = combined[l]
- important_gene_dataframe.head()
- # %%
- import pandas as pd
- import seaborn as sns
- import matplotlib.pyplot as plt
- from sklearn.preprocessing import MinMaxScaler
- # Set the Seaborn theme for better aesthetics
- sns.set_theme(style="whitegrid", palette="muted")
- # Assuming 'important_gene_dataframe' is your DataFrame and it has a 'Label' column
- # Normalize the entire DataFrame (excluding the 'Label' column) to a 0-1 range
- scaler = MinMaxScaler()
- important_gene_dataframe_normalized = important_gene_dataframe.copy()
- # Apply normalization (only to the gene expression columns, not the 'Label' column)
- gene_columns = important_gene_dataframe_normalized.drop(columns=['Label']).columns
- important_gene_dataframe_normalized[gene_columns] = scaler.fit_transform(important_gene_dataframe_normalized[gene_columns])
- # Filter DataFrames for 'AD' and 'NCI'
- ad_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'AD']
- nci_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'NCI']
- # Drop the 'Label' column and only keep the gene expression values
- ad_data = ad_data.drop(columns=['Label'])
- nci_data = nci_data.drop(columns=['Label'])
- # List of genes (columns) to iterate over
- genes = ad_data.columns
- # Number of genes
- num_genes = len(genes)
- # Determine the number of rows and columns for subplots (depending on the number of genes)
- ncols = 12 # Set the number of columns (adjust based on your preference)
- nrows = (num_genes + ncols - 1) // ncols # Automatically adjust the number of rows
- # Create a figure and axes for subplots
- fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(50, 10 * nrows))
- axes = axes.flatten() # Flatten axes array for easy indexing
- # Loop through each gene and create a boxplot on a subplot
- for idx, gene in enumerate(genes):
- ax = axes[idx]
- # Prepare the data for the current gene (AD vs NCI)
- data = pd.DataFrame({
- 'Group': ['AD'] * len(ad_data[gene]) + ['NCI'] * len(nci_data[gene]),
- 'Expression': list(ad_data[gene]) + list(nci_data[gene])
- })
- # Create a boxplot for the current gene
- sns.boxplot(data=data, x='Group', y='Expression', palette='pastel', ax=ax)
- gene = n[idx]
- # Add a title and custom labels
- ax.set_title(f'{gene}', fontsize=14, weight='bold')
- ax.set_xlabel('Group', fontsize=12)
- ax.set_ylabel('Normalized Expression (0-1)', fontsize=12)
- # Add gridlines for better readability
- ax.grid(True, axis='y', linestyle='--', alpha=0.6)
- # Remove any empty subplots if the number of genes is not a perfect multiple of ncols
- for idx in range(num_genes, len(axes)):
- fig.delaxes(axes[idx])
- # Adjust layout for better spacing between subplots
- plt.tight_layout()
- # Save the figure as an SVG with high resolution
- fig.savefig('gene_expression_boxplots.svg', format='svg', dpi=600)
- # Show the plot
- plt.show()
- # %%
- l.remove('Label')
- plt.figure(figsize=(8, 6)) # Optional: adjust the figure size
- sns.heatmap(important_gene_dataframe[l], annot=False, cmap='YlGnBu') # 'annot=True' adds annotations inside the heatmap cells
- plt.title('Heatmap of DataFrame')
- plt.show()
- # %%
- AD_panda = pd.read_csv('../../../data_preprocessed/grn_output/ad_nci_seperate_network/panda/output_panda_ad_csv.csv')
- # %%
- AD_panda.sort_values(by = 'force', ascending=False)
- # %%
- import pandas as pd
- import seaborn as sns
- import matplotlib.pyplot as plt
- import numpy as np
- # Set the plot size
- plt.figure(figsize=(8, 6))
- # Plot percentile plot using seaborn
- sns.ecdfplot(AD_panda['force'], color='skyblue')
- percentile_95 = np.percentile(AD_panda['force'], 95)
- plt.axvline(x=percentile_95, color='red', linestyle='--', label='90th Percentile')
- plt.axhline(y=0.9, color='green', linestyle='--')
- # Annotate the point on the plot
- plt.annotate(f'({percentile_95:.2f}, 0.95)', xy=(percentile_95, 0.9), xytext=(percentile_95 + 0.1, 0.9),
- arrowprops=dict(facecolor='black', arrowstyle='->'))
- # Display the plot
- plt.title('Percentile Plot (CDF)')
- plt.xlabel('Data')
- plt.ylabel('Cumulative Probability')
- plt.show()
- # %%
- AD_panda_top_5 = AD_panda[AD_panda.force>=3.71]
- AD_panda_top_5
- # %%
- list_of_tf = list(set(AD_panda_top_5.tf))
- AD_panda_top_5_bipartite_complete = AD_panda_top_5[~AD_panda_top_5.gene.isin(list_of_tf)]
- # %%
- import pandas as pd
- import networkx as nx
- import matplotlib.pyplot as plt
- from tqdm import tqdm
- # Create a weighted bipartite graph
- AD_BG_bipartite = nx.Graph()
- AD_BG_bipartite.add_nodes_from(AD_panda_top_5_bipartite_complete['tf'], bipartite = 0)
- AD_BG_bipartite.add_nodes_from(AD_panda_top_5_bipartite_complete['gene'], bipartite = 1)
- # Add weighted edges
- edges = [(row['tf'], row['gene'], {'weight': row['force']}) for idx, row in tqdm(AD_panda_top_5_bipartite_complete.iterrows())]
- AD_BG_bipartite.add_edges_from(edges)
- # %%
- relevant_genes = l.copy()
- # %%
- relevant_genes
- # %%
- l = []
- for i in relevant_genes:
- l.append(i.split('_')[0])
- l.append(i.split('_')[1])
- # %%
- import gseapy as gp
- from gseapy import Msigdb
- names = gp.get_library_name(organism='Human')
- # %%
- l[0:5]
- # %%
- def return_gene_name(gene):
- try:
- if ensemble_to_gene_name[gene] == 'Gene name not found':
- return gene
- return ensemble_to_gene_name[gene]
- except:
- return gene
- # %%
- 0%2 != 0
- # %%
- 2%2 != 0
- # %%
- df = pd.DataFrame()
- data_dict = dict()
- for i,k in tqdm(enumerate(l)):
- if i%2 != 0:
- continue
- m = []
- try:
- #print(k)
- #print(l[i+1])
- m.append(k)
- m.extend(list(AD_BG_bipartite[k]))
- m.extend(list(AD_BG_bipartite[l[i+1]]))
- except:
- print(k)
- m = set(m)
- print(len(m))
- gene_set = [return_gene_name(gene_id) for gene_id in m]
- gene_set = list(set(gene_set))
- try:
- gene_set.remove('Gene name not found')
- except:
- print('skip')
- filtered_gene_set = []
- k = return_gene_name(k)
- h = return_gene_name(l[i+1])
- for p,j in enumerate(gene_set):
- if type(j) == type(gene_set[0]):
- filtered_gene_set.append(j)
- else:
- continue
- print(len(filtered_gene_set))
- enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
- df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
- df2 = df2[df2['Adjusted P-value']<0.05]
- data_dict[f'{k}-{h}'] = list(df2.Term)
- print(f'{k}-{h}')
- df = pd.concat([df,df2])
- # %%
- df
- # %%
- data_dict.keys()
- # %%
- import pandas as pd
- from tabulate import tabulate
- from openpyxl import Workbook, load_workbook
- from openpyxl.styles import Border, Side, Font, PatternFill
- # Input dictionary
- data = data_dict
- # Extract unique pathways
- pathways = sorted(set(p for v in data.values() for p in v))
- columns = list(data.keys())
- df = pd.DataFrame(index=pathways, columns=columns)
- column_widths = {
- "A": 80, # Increase column A width
- "B": 25, # Increase column B width
- "C": 25, # Increase column C width
- "D": 25,
- "E":25,
- "F": 25, # Increase column A width
- "G": 25, # Increase column B width
- "H": 25, # Increase column C width
- "I": 25,
- "K":25,
- "L": 25, # Increase column A width
- "M": 25, # Increase column B width
- "N": 25, # Increase column C width
- "O": 25,
- "P":25,
- "Q": 25, # Increase column A width
- "R": 25, # Increase column B width
- "S": 25, # Increase column C width
- "T": 25,
- "U":25,
- "V": 25, # Increase column A width
- "W": 25 #Increase column D width
- }
- # Function to insert tick mark
- def tick():
- return "✔"
- # Fill DataFrame with tick marks
- for key, values in data.items():
- for value in values:
- df.loc[value, key] = tick()
- df.fillna("", inplace=True)
- # Define styles
- thin_border = Border(left=Side(style='thin'),
- right=Side(style='thin'),
- top=Side(style='thin'),
- bottom=Side(style='thin'))
- bold_font = Font(bold=True)
- gray_fill = PatternFill(start_color="D3D3D3", end_color="D3D3D3", fill_type="solid")
- # Set chunk sizes
- columns_per_sheet = 25
- rows_per_sheet = 10000
- # Split into chunks of 8 columns and 28 rows
- total_col_chunks = (len(df.columns) // columns_per_sheet) + (1 if len(df.columns) % columns_per_sheet else 0)
- total_row_chunks = (len(df) // rows_per_sheet) + (1 if len(df) % rows_per_sheet else 0)
- file_count = 1
- for col_chunk in range(total_col_chunks):
- for row_chunk in range(total_row_chunks):
- start_col = col_chunk * columns_per_sheet
- end_col = start_col + columns_per_sheet
- start_row = row_chunk * rows_per_sheet
- end_row = start_row + rows_per_sheet
- df_chunk = df.iloc[start_row:end_row, start_col:end_col]
- # Ensure the chunk has the correct dimensions
- while len(df_chunk.columns) < columns_per_sheet:
- df_chunk[f"Extra_{len(df_chunk.columns)+1}"] = ""
- excel_path = f"/12tb_dsk2/danish/GRN_new_21_feb_2025/data_preprocessed/grn_analysis_output/gsea_excel_sheets/pathway_table_part_{file_count}.xlsx"
- df_chunk.to_excel(excel_path)
- # Load the workbook and apply formatting
- wb = load_workbook(excel_path)
- ws = wb.active
- # Freeze headers
- ws.freeze_panes = "B2" # Freeze first row and column
- # Apply styles
- for row in ws.iter_rows():
- for cell in row:
- cell.border = thin_border
- if cell.value == "✔":
- cell.font = bold_font
- cell.fill = gray_fill
- # Increase row height
- for row in ws.iter_rows():
- ws.row_dimensions[row[0].row].height = 20
- for col_letter, width in column_widths.items():
- ws.column_dimensions[col_letter].width = width
- wb.save(excel_path)
- file_count += 1
- print("Excel files created successfully! ✅")
- # %%
- df
- # %%
- # Convert DataFrame to HTML
- html_path = "temp.html"
- df.to_html(html_path, index=False)
- # Convert HTML to PDF
- pdf_path = "output.pdf"
- pdfkit.from_file(html_path, pdf_path)
- print(f"Excel file converted to PDF at {pdf_path}")
- # %%
- # Convert DataFrame to HTML
- html_path = "temp.html"
- df.to_html(html_path, index=False)
- # Convert HTML to PDF
- pdf_path = "output.pdf"
- pdfkit.from_file(html_path, pdf_path)
- print(f"Excel file converted to PDF at {pdf_path}")
- # %%
- hub_tf= ['ENSG00000186300','ENSG00000256294','ENSG00000234284','ENSG00000188785']
- # %%
- hub_tf = set(l).intersection(set(hub_tf))
- # %%
- hub_tf = list(hub_tf)
- hub_tf
- # %%
- relevant_genes
- # %%
- hub_tf = ['ENSG00000188785', 'ENSG00000234284', 'ENSG00000256294']
- # %%
- hub_attached_genes = ['ENSG00000122641','ENSG00000149089','ENSG00000272576'] # manually checked
- # %%
- m = []
- for k,i in enumerate(hub_tf):
- print(return_gene_name(hub_tf[k]))
- print(return_gene_name(hub_attached_genes[k]))
- m.append(f'{return_gene_name(hub_tf[k])}-{return_gene_name(hub_attached_genes[k])}')
- # %%
- m
- # %%
- m.append('Label')
- # %%
- n
- # %%
- n.append('Label')
- n
- # %%
- important_gene_dataframe.columns = n
- # %%
- important_gene_dataframe
- # %%
- df = important_gene_dataframe[m].copy()
- # Splitting data by labels
- df
- # %%
- important_gene_dataframe.columns
- # %%
- ad_data = df[df['Label'] == 'AD'].drop(columns='Label')
- nci_data = df[df['Label'] == 'NCI'].drop(columns='Label')
- from scipy import stats
- # Initialize dictionary to store p-values for each column
- p_values = {}
- # Perform median test for each column
- for column in ad_data.columns:
- # Get the data for the current column in each group
- group1 = ad_data[column]
- group2 = nci_data[column]
- # Perform the median test
- stat, p_value, med, table = stats.median_test(group1, group2)
- # Store the p-value for the column
- p_values[column] = p_value
- # Output the p-values
- print(p_values)
- # %%
- ad_data[['ZNF548-INHBA','ZNF879-APIP']]
- # %%
- # %%
- import pandas as pd
- import seaborn as sns
- import matplotlib.pyplot as plt
- from sklearn.preprocessing import MinMaxScaler
- # Set the Seaborn theme for better aesthetics
- sns.set_theme(style="whitegrid", palette="muted")
- # Assuming 'important_gene_dataframe' is your DataFrame and it has a 'Label' column
- # Normalize the entire DataFrame (excluding the 'Label' column) to a 0-1 range
- scaler = MinMaxScaler()
- important_gene_dataframe_normalized = important_gene_dataframe[m].copy()
- # Apply normalization (only to the gene expression columns, not the 'Label' column)
- gene_columns = important_gene_dataframe_normalized.drop(columns=['Label']).columns
- important_gene_dataframe_normalized[gene_columns] = scaler.fit_transform(important_gene_dataframe_normalized[gene_columns])
- # Filter DataFrames for 'AD' and 'NCI'
- ad_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'AD']
- nci_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'NCI']
- # Drop the 'Label' column and only keep the gene expression values
- ad_data = ad_data.drop(columns=['Label'])
- nci_data = nci_data.drop(columns=['Label'])
- # List of genes (columns) to iterate over
- genes = ad_data.columns
- # Number of genes
- num_genes = len(genes)
- ad_data = remove_outliers(ad_data)
- nci_data = remove_outliers(nci_data)
- # Determine the number of rows and columns for subplots (depending on the number of genes)
- ncols = 12 # Set the number of columns (adjust based on your preference)
- nrows = (num_genes + ncols - 1) // ncols # Automatically adjust the number of rows
- # Create a figure and axes for subplots
- fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(50, 10 * nrows))
- axes = axes.flatten() # Flatten axes array for easy indexing
- # Loop through each gene and create a boxplot on a subplot
- for idx, gene in enumerate(genes):
- ax = axes[idx]
- # Prepare the data for the current gene (AD vs NCI)
- data = pd.DataFrame({
- 'Group': ['AD'] * len(ad_data[gene]) + ['NCI'] * len(nci_data[gene]),
- 'Expression': list(ad_data[gene]) + list(nci_data[gene])
- })
- # Create a boxplot for the current gene
- sns.boxplot(data=data, x='Group', y='Expression', palette='pastel', ax=ax)
- gene = m[idx]
- # Add a title and custom labels
- ax.set_title(f'{gene}', fontsize=14, weight='bold')
- ax.set_xlabel('Group', fontsize=12)
- ax.set_ylabel('Normalized Expression (0-1)', fontsize=12)
- # Add gridlines for better readability
- max_y = max(data['Expression']) + 0.05
- line_y = max_y + 0.02 # Line position
- cap_height = 0.01 # Height of vertical caps
- # Add horizontal line for significance comparison
- ax.hlines(y=line_y, xmin=0, xmax=1, color='black', linewidth=1.5)
- # Add vertical caps at both ends of the horizontal line
- ax.vlines(x=0, ymin=line_y - cap_height, ymax=line_y, color='black', linewidth=1.5)
- ax.vlines(x=1, ymin=line_y - cap_height, ymax=line_y, color='black', linewidth=1.5)
- # Display p-value above the plot
- ax.text(0.5, line_y + 0.02, f"p = {p_value:.3f}", ha='center', va='bottom', fontsize=12, color='black')
- # Gridlines for better readability
- ax.grid(True, axis='y', linestyle='--', alpha=0.6)
- # Remove any empty subplots if the number of genes is not a perfect multiple of ncols
- for idx in range(num_genes, len(axes)):
- fig.delaxes(axes[idx])
- # Adjust layout for better spacing between subplots
- plt.tight_layout()
- # Save the figure as an SVG with high resolution
- fig.savefig('/12tb_dsk2/danish/GRN_new_21_feb_2025/data_preprocessed/grn_analysis_output/plots/gene_expression_boxplots_subset_after_removing_outliers.svg', format='svg', dpi=300)
- # Show the plot
- plt.show()
- # %%
- import pandas as pd
- import seaborn as sns
- import matplotlib.pyplot as plt
- from sklearn.preprocessing import MinMaxScaler
- # Set the Seaborn theme for better aesthetics
- sns.set_theme(style="whitegrid", palette="muted")
- # Assuming 'important_gene_dataframe' is your DataFrame and it has a 'Label' column
- scaler = MinMaxScaler()
- # Copy the DataFrame and exclude the 'Label' column from normalization
- important_gene_dataframe_normalized = important_gene_dataframe.copy()
- gene_columns = important_gene_dataframe.drop(columns=['Label']).columns
- important_gene_dataframe_normalized[gene_columns] = scaler.fit_transform(important_gene_dataframe[gene_columns])
- # Filter DataFrames for 'AD' and 'NCI'
- ad_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'AD'].copy()
- nci_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'NCI'].copy()
- # Remove the 'Label' column to work only with gene expression values
- ad_data = ad_data.drop(columns=['Label'])
- nci_data = nci_data.drop(columns=['Label'])
- # Function to remove outliers using IQR
- def remove_outliers(df):
- Q1 = df.quantile(0.25) # First quartile
- Q3 = df.quantile(0.75) # Third quartile
- IQR = Q3 - Q1 # Interquartile range
- lower_bound = Q1 - 1.5 * IQR
- upper_bound = Q3 + 1.5 * IQR
- return df[~((df < lower_bound) | (df > upper_bound)).any(axis=1)]
- # Apply outlier removal
- ad_data_clean = remove_outliers(ad_data)
- nci_data_clean = remove_outliers(nci_data)
- # List of genes to iterate over
- genes = ad_data_clean.columns
- num_genes = len(genes)
- # Determine subplot layout
- ncols = 12
- nrows = (num_genes + ncols - 1) // ncols # Auto-adjust rows
- # Create figure
- fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(50, 10 * nrows))
- axes = axes.flatten()
- # Loop through each gene to create boxplots
- for idx, gene in enumerate(genes):
- ax = axes[idx]
- # Prepare data for plotting
- data = pd.DataFrame({
- 'Group': ['AD'] * len(ad_data_clean[gene]) + ['NCI'] * len(nci_data_clean[gene]),
- 'Expression': list(ad_data_clean[gene]) + list(nci_data_clean[gene])
- })
- # Create boxplot
- sns.boxplot(data=data, x='Group', y='Expression', palette='pastel', ax=ax)
- # Set title and labels
- ax.set_title(f'{gene}', fontsize=14, weight='bold')
- ax.set_xlabel('Group', fontsize=12)
- ax.set_ylabel('Normalized Expression (0-1)', fontsize=12)
- # Grid for better readability
- ax.grid(True, axis='y', linestyle='--', alpha=0.6)
- # Remove empty subplots
- for idx in range(num_genes, len(axes)):
- fig.delaxes(axes[idx])
- # Adjust layout
- plt.tight_layout()
- # Show plot
- plt.show()
- # %%
- m
- # %%
- df = pd.DataFrame()
- data_dict = dict()
- for i,k in tqdm(enumerate(l)):
- if i%2 != 0:
- continue
- m = []
- try:
- #print(k)
- #print(l[i+1])
- m.append(k)
- m.extend(list(AD_BG_bipartite[k]))
- m.extend(list(AD_BG_bipartite[l[i+1]]))
- except:
- print(k)
- m = set(m)
- print(len(m))
- gene_set = [return_gene_name(gene_id) for gene_id in m]
- gene_set = list(set(gene_set))
- try:
- gene_set.remove('Gene name not found')
- except:
- print('skip')
- filtered_gene_set = []
- k = return_gene_name(k)
- h = return_gene_name(l[i+1])
- for p,j in enumerate(gene_set):
- if type(j) == type(gene_set[0]):
- filtered_gene_set.append(j)
- else:
- continue
- print(len(filtered_gene_set))
- enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
- df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
- df2 = df2[df2['Adjusted P-value']<0.005]
- data_dict[f'{k}-{h}'] = list(df2.Term)
- print(f'{k}-{h}')
- df = pd.concat([df,df2])
- # %%
- # 'GO_Biological_Process_2023'
- # %%
- df.Gene_set.value_counts()
- # %%
- df
- # %%
- import plotly.graph_objects as go
- # Data preparation as before
- sources = []
- targets = []
- for gene, pathways in data_dict.items():
- for pathway in pathways:
- sources.append(gene)
- targets.append(pathway)
- # Create a mapping of genes and pathways to unique indices
- unique_items = list(set(sources + targets))
- item_indices = {item: idx for idx, item in enumerate(unique_items)}
- # Create lists of indices for the Sankey diagram
- source_indices = [item_indices[gene] for gene in sources]
- target_indices = [item_indices[pathway] for pathway in targets]
- # Define the Sankey diagram figure
- fig = go.Figure(go.Sankey(
- node=dict(
- pad=15,
- thickness=20,
- line=dict(color="black", width=0.5),
- label=unique_items,
- align='right'
- ),
- link=dict(
- source=source_indices,
- target=target_indices,
- value=[1] * len(sources), # All links have equal value
- )
- ))
- # Update layout with increased figure size
- fig.update_layout(
- title_text="Gene to Pathway Sankey Diagram",
- font_size=12,
- width=900, # Increase width
- height=1550, # Increase height
- )
- # Save the figure as SVG
- fig.write_image("/12tb_dsk2/danish/GRN_new_21_feb_2025/data_preprocessed/grn_analysis_output/plots/sankey_enriched_plots.svg", format="svg", scale=1)# Show the figure
- fig.show()
- # %%
- set(df.Term)
- # %%
- df = pd.DataFrame()
- data_dict = dict()
- for i,k in tqdm(enumerate(l)):
- if i%2 != 0:
- continue
- m = []
- try:
- #print(k)
- #print(l[i+1])
- m.append(k)
- m.extend(list(AD_BG_bipartite[k]))
- m.extend(list(AD_BG_bipartite[l[i+1]]))
- except:
- print(k)
- m = set(m)
- print(len(m))
- gene_set = [return_gene_name(gene_id) for gene_id in m]
- gene_set = list(set(gene_set))
- try:
- gene_set.remove('Gene name not found')
- except:
- print('skip')
- filtered_gene_set = []
- k = return_gene_name(k)
- h = return_gene_name(l[i+1])
- for p,j in enumerate(gene_set):
- if type(j) == type(gene_set[0]):
- filtered_gene_set.append(j)
- else:
- continue
- print(len(filtered_gene_set))
- enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
- df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
- df2 = df2[df2['Adjusted P-value']<0.05]
- data_dict[f'{k}-{h}'] = list(df2.Term)
- print(f'{k}-{h}')
- df = pd.concat([df,df2])
- # %%
- data_dict.keys()
- # %%
- data_dict['MTF1-CYLD']
- # %%
- relevant_genes
- # %%
- AD_related_genes = pd.read_csv('../../../data_preprocessed/previously_researched_data/gene-list.csv')
- AD_related_genes
- # %%
- gene_name_to_ensemble = {v: k for k, v in ensemble_to_gene_name.items()}
- # %%
- def return_ensemble_name(gene_name):
- try:
- return gene_name_to_ensemble[gene_name]
- except:
- return gene_name
- # %%
- AD_related_genes_list = [return_ensemble_name(gene_name) for gene_name in AD_related_genes['Gene Symbol'].tolist()]
- # %%
- AD_related_genes_list[0:10]
- # %%
- s = set()
- for n in range(0,len(l),2):
- k = l[n]
- connected_nodes = list(AD_BG_bipartite.neighbors(k))
- genes_l = set(connected_nodes).intersection(set(AD_related_genes_list))
- x = len(set(connected_nodes).intersection(set(AD_related_genes_list)))
- if x>0:
- print(k)
- print(len(connected_nodes))
- print(len(genes_l))
- print(len(genes_l)/len(connected_nodes))
- print(genes_l)
- s.update(genes_l)
- # %%
- len(s)
- # %%
- len(s)/948
- # %%
- import networkx as nx
- # Assuming l is your list of nodes and AD_BG_bipartite is your bipartite graph
- # Initialize a set to store the nodes to be included in the subgraph
- nodes_to_include = set(l)
- for n in range(0,len(l),2):
- k = l[n]
- connected_nodes = list(AD_BG_bipartite.neighbors(k))
- nodes_to_include.update(connected_nodes)
- # Now create the subgraph containing only the nodes in nodes_to_include
- subgraph = AD_BG_bipartite.subgraph(nodes_to_include).copy()
- # subgraph now contains the nodes in 'l' and their neighbors
- # %%
- highlighted_edges = []
- for i in relevant_genes:
- tf = i.split('_')[0]
- gene = i.split('_')[1]
- highlighted_edges.append((tf, gene))
- # %%
- import networkx as nx
- import matplotlib.pyplot as plt
- from adjustText import adjust_text
- # Assuming l is your list of nodes and AD_BG_bipartite is your bipartite graph
- # Initialize a set to store the nodes to be included in the subgraph
- nodes_to_include = set()
- s = []
- # Add neighbors to nodes_to_include
- for n in range(0, len(l), 2):
- k = l[n]
- nodes_to_include.update(set([k]))
- s.append(l[n])
- connected_nodes = list(AD_BG_bipartite.neighbors(k))
- genes_l = set(connected_nodes).intersection(set(AD_related_genes_list))
- genes_l.update(l)
- nodes_to_include.update(genes_l)
- # Now create the subgraph containing only the nodes in nodes_to_include
- subgraph = AD_BG_bipartite.subgraph(nodes_to_include).copy()
- # Position nodes using a layout (for example, spring layout)
- pos = nx.spring_layout(subgraph)
- # Draw nodes and edges
- plt.figure(figsize=(20, 20))
- node_sizes = []
- # Draw the nodes
- # Nodes from list 'l' will be squares, others will be circles
- node_shapes = []
- for node in subgraph.nodes():
- if node in s:
- node_shapes.append('s')
- node_sizes.append(380) # Square for nodes in 'l'
- else:
- node_shapes.append('o')
- node_sizes.append(250) # Circle for connected nodes
- # Draw the graph with node shapes based on 'l' membership
- nx.draw_networkx_edges(subgraph, pos, width=1.0, alpha=0.7, edge_color='blue')
- # Draw the nodes with different shapes (squares for 'l' and circles for others)
- for i, node in enumerate(subgraph.nodes()):
- shape = node_shapes[i]
- size = node_sizes[i]
- nx.draw_networkx_nodes(subgraph, pos, nodelist=[node], node_size=size, node_color='blue', node_shape=shape)
- nx.draw_networkx_nodes(subgraph, pos, nodelist=s, node_size=size, node_color='red',node_shape=shape)
- gene_imp = set(l) - set(s)
- gene_imp = list(gene_imp)
- nx.draw_networkx_nodes(subgraph, pos, nodelist=gene_imp, node_size=450, node_color='red',node_shape='o')
- nx.draw_networkx_edges(subgraph, pos, edgelist=highlighted_edges, width=3, alpha=0.8, edge_color='red')
- # Labeling (only labels for nodes in 'l')
- labels = {node: return_gene_name(node) for node in l} # Label nodes from 'l' only
- texts = []
- for node, label in labels.items():
- text = plt.text(pos[node][0], pos[node][1], label, fontsize=6, fontweight='bold', color='black', ha='left')
- texts.append(text)
- # Adjust label positions to avoid overlap
- adjust_text(texts, arrowprops=dict(arrowstyle="->", color='black', lw=1))
- # Finalize plot
- plt.title("Subgraph with Adjusted Labels for 'l' Nodes", fontsize=16)
- plt.axis('off') # Remove axis
- plt.savefig('../../../data_preprocessed/grn_analysis_output/plots/network.svg', format='svg', dpi=600) # Corrected .savefig -> plt.savefig
- plt.show()
- # %%
- labels
- # %% [markdown]
- # # AD-targeted Genes-collected: https://agora.adknowledgeportal.org/genes/nominated-targets
- # %%
- # %%
- # %%
- # %%
- # %%
- # %%
- # %%
- # %%
- # %%
- # %%
- # %%
- TNF-alpha signaling via NF-kB, structural components like collagen-containing extracellular matrix (GO:0062023) and microvillus (GO:0005902), as well as molecular activities such as hormone activity (GO:0005179)
- # %%
- 'collagen-containing extracellular matrix (GO:0062023)' in df.Term.tolist()
- # %%
- df.Term.tolist()
- # %%
- df = pd.DataFrame()
- data_dict = dict()
- for i,k in tqdm(enumerate(l)):
- if i%2 != 0:
- continue
- m = []
- try:
- #print(k)
- #print(l[i+1])
- m.append(k)
- m.extend(list(AD_BG_bipartite[k]))
- m.extend(list(AD_BG_bipartite[l[i+1]]))
- except:
- print(k)
- m = set(m)
- print(len(m))
- gene_set = [return_gene_name(gene_id) for gene_id in m]
- gene_set = list(set(gene_set))
- try:
- gene_set.remove('Gene name not found')
- except:
- print('skip')
- filtered_gene_set = []
- k = return_gene_name(k)
- h = return_gene_name(l[i+1])
- for p,j in enumerate(gene_set):
- if type(j) == type(gene_set[0]):
- filtered_gene_set.append(j)
- else:
- continue
- print(len(filtered_gene_set))
- enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
- df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
- df2 = df2[df2['Adjusted P-value']<0.05]
- data_dict[f'{k}-{h}'] = list(df2.Term)
- print(f'{k}-{h}')
- df = pd.concat([df,df2])
- # %%
- df
- # %%
- 'collagen-containing extracellular matrix (GO:0062023)' in df.Term.tolist()
- # %%
- diff_exp_enriched = set(['TNF-alpha Signaling via NF-kB', 'Collagen-Containing Extracellular Matrix (GO:0062023)','Microvillus (GO:0005902)', 'Hormone Activity (GO:0005179)', 'Neutral L-amino Acid:Sodium Symporter Activity (GO:0005295)'])
- # %%
- diff_exp_enriched - set(df.Term.tolist())
- # %%
- diff_exp_enriched.
- # %%
- df = pd.DataFrame()
- data_dict = dict()
- for i,k in tqdm(enumerate(l)):
- if i%2 != 0:
- continue
- m = []
- try:
- #print(k)
- #print(l[i+1])
- m.append(k)
- m.extend(list(AD_BG_bipartite[k]))
- m.extend(list(AD_BG_bipartite[l[i+1]]))
- except:
- print(k)
- m = set(m)
- print(len(m))
- gene_set = [return_gene_name(gene_id) for gene_id in m]
- gene_set = list(set(gene_set))
- try:
- gene_set.remove('Gene name not found')
- except:
- print('skip')
- filtered_gene_set = []
- k = return_gene_name(k)
- h = return_gene_name(l[i+1])
- for p,j in enumerate(gene_set):
- if type(j) == type(gene_set[0]):
- filtered_gene_set.append(j)
- else:
- continue
- print(len(filtered_gene_set))
- enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
- df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
- df2 = df2[df2['Adjusted P-value']<0.005]
- data_dict[f'{k}-{h}'] = list(df2.Term)
- print(f'{k}-{h}')
- df = pd.concat([df,df2])
- # %%
- diff_exp_enriched - set(df.Term.tolist())
- # %%
- # %%
- # %%
plots(box_plot_sankey_plots_network_plots).ipynb at commit 3f3f8d9, no license · at the source
Overview
- Division of Systems and Synthetic Biology, Department of Life Sciences, Chalmers University of Technology,Gothenburg, Sweden
- Department of Microbiology, Oslo University Hospital and University of Oslo,Oslo, Norway
Abstract
Alzheimer’s disease (AD) is a multifactorial neurodegenerative disorder marked by progressive cognitive decline, yet its transcriptional regulatory architecture remains poorly understood. Here, we model sample-specific gene regulatory networks (GRNs) from dorsolateral prefrontal cortex transcriptomes of 87 individuals with AD and 67 non-cognitively impaired (NCI) controls and use a machine learning classifier to detect consistent disease-specific network features. This sample-specific network approach captures inter-individual variation in transcriptional regulation and revealed 22 key transcription factor-gene regulations that distinguish AD from NCI with 96% weighted accuracy. The key transcription factor-gene interactions were enriched in pathways central to AD pathology, including synaptic signalling, mitochondrial function, proteostasis, and neuroinflammation. Network analysis uncovered significant differences in regulatory connectivity between AD and controls, with ZNF225, ZNF849, and ZNF548 emerging as AD-specific regulatory hubs. Moreover, several key regulatory edges showed significant correlations with longitudinal cognitive decline, supporting their clinical relevance. Our findings highlight pervasive transcriptional dysregulation in AD, emphasizing sample-specific GRN modelling’s value in uncovering regulatory mechanisms.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.
Polster-lab/Analysis_GRN_ROSMAP
3f3f8d9e1bb54c67b99042add4a411fbae4b0127, 3 April 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
39 files
- data_preprocessing.ipynb
, Jupyter, 1,539 lines - src/
code_for_aggregate_netwo , Jupyter, 269 linesrk_panda/ .ipynb_checkpoints/ motif_ppi_preprocessing_ and_panda-checkpoint.ipy nb - src/
code_for_aggregate_netwo , Jupyter, 73 linesrk_panda/ .ipynb_checkpoints/ panda_lioness_ad-checkpo int.ipynb - src/
code_for_aggregate_netwo , Jupyter, 31 linesrk_panda/ .ipynb_checkpoints/ panda_lioness_ad_nci_tog ether-checkpoint.ipynb - src/
code_for_aggregate_netwo , Jupyter, 73 linesrk_panda/ .ipynb_checkpoints/ panda_lioness_nci-checkp oint.ipynb - src/
code_for_aggregate_netwo , Jupyter, 73 linesrk_panda/ panda_lioness_ad.ipynb - src/
code_for_aggregate_netwo , Jupyter, 67 linesrk_panda/ panda_lioness_ad_nci_tog ether.ipynb - src/
code_for_aggregate_netwo , Jupyter, 73 linesrk_panda/ panda_lioness_nci.ipynb - src/
code_for_data_preprocess , Jupyter, 1 lineing/ .ipynb_checkpoints/ clinical_data_preprocess ing-checkpoint.ipynb - src/
code_for_data_preprocess , Jupyter, 1 lineing/ .ipynb_checkpoints/ combat_seq_batch_removal -checkpoint.ipynb - src/
code_for_data_preprocess , Jupyter, 1,536 linesing/ .ipynb_checkpoints/ data_preprocessing-check point.ipynb - src/
code_for_data_preprocess , Jupyter, 269 linesing/ .ipynb_checkpoints/ motif_ppi_preprocessing- checkpoint.ipynb - src/
code_for_data_preprocess , Jupyter, 361 linesing/ .ipynb_checkpoints/ motif_ppi_preprocessing_ and_panda-checkpoint.ipy nb - src/
code_for_data_preprocess , Jupyter, 363 linesing/ clinical_data_preprocess ing.ipynb - src/
code_for_data_preprocess , Jupyter, 81 lines, 1 matching/ combat_seq_batch_removal .ipynb - src/
code_for_data_preprocess , Jupyter, 1,566 linesing/ data_preprocessing.ipynb - src/
code_for_data_preprocess , Jupyter, 361 linesing/ motif_ppi_preprocessing_ and_panda.ipynb - src/
code_for_lioness_network , Jupyter, 69 liness_analysis/ .ipynb_checkpoints/ Agg_GRN_seperate_ad_nci_ analysis-checkpoint.ipyn b - src/
code_for_lioness_network , Jupyter, 134 liness_analysis/ .ipynb_checkpoints/ Agg_GRN_together_ad_nci_ analysis-checkpoint.ipyn b - src/
code_for_lioness_network , Jupyter, 1,069 liness_analysis/ .ipynb_checkpoints/ sample_specific_analysis _ad_nci_seperate-checkpo int.ipynb - src/
code_for_lioness_network , Jupyter, 82 liness_analysis/ .ipynb_checkpoints/ sample_specific_analysis _ad_nci_together-checkpo int.ipynb - src/
code_for_lioness_network , Jupyter, 134 liness_analysis/ Agg_GRN_seperate_ad_nci_ analysis.ipynb - src/
code_for_lioness_network , Jupyter, 72 liness_analysis/ Agg_GRN_together_ad_nci_ analysis.ipynb - src/
code_for_lioness_network , Jupyter, 1,283 liness_analysis/ Sankey_Plot/ .ipynb_checkpoints/ plots(box_plot_sankey_pl ots_network_plots)-check point.ipynb - src/
code_for_lioness_network , Jupyter, 1,329 lines, 2 matchess_analysis/ Sankey_Plot/ plots(box_plot_sankey_pl ots_network_plots).ipynb - src/
code_for_lioness_network , Jupyter, 1 lines_analysis/ calculating_graph_featur es/ .ipynb_checkpoints/ ad_netwoks_features_calc ulation-checkpoint.ipynb - src/
code_for_lioness_network , Jupyter, 127 liness_analysis/ calculating_graph_featur es/ .ipynb_checkpoints/ nci_netwoks_features_cal culation-checkpoint.ipyn b - src/
code_for_lioness_network , Jupyter, 141 liness_analysis/ calculating_graph_featur es/ .ipynb_checkpoints/ netwoks_features_calcula tion-checkpoint.ipynb - src/
code_for_lioness_network , Jupyter, 137 liness_analysis/ calculating_graph_featur es/ .ipynb_checkpoints/ output_analysis-checkpoi nt.ipynb - src/
code_for_lioness_network , Jupyter, 129 lines, 2 matchess_analysis/ calculating_graph_featur es/ ad_netwoks_features_calc ulation.ipynb - src/
code_for_lioness_network , Jupyter, 112 lines, 2 matchess_analysis/ calculating_graph_featur es/ nci_netwoks_features_cal culation.ipynb - src/
code_for_lioness_network , Jupyter, 141 liness_analysis/ calculating_graph_featur es/ netwoks_features_calcula tion.ipynb - src/
code_for_lioness_network , Jupyter, 169 liness_analysis/ calculating_graph_featur es/ output_analysis.ipynb - src/
code_for_lioness_network , Jupyter, 1 lines_analysis/ cell_markers_check/ .ipynb_checkpoints/ cell_markers_check-check point.ipynb - src/
code_for_lioness_network , Jupyter, 172 liness_analysis/ cell_markers_check/ cell_markers_check.ipynb - src/
code_for_lioness_network , Jupyter, 1 lines_analysis/ correlation_analysis/ .ipynb_checkpoints/ bubble_plots-checkpoint. ipynb - src/
code_for_lioness_network , Jupyter, 1 lines_analysis/ correlation_analysis/ .ipynb_checkpoints/ comparing_correlations-c heckpoint.ipynb - src/
code_for_lioness_network , Jupyter, 299 liness_analysis/ correlation_analysis/ .ipynb_checkpoints/ correlation_anaysis-chec kpoint.ipynb - src/
code_for_lioness_network , Jupyter, 637 liness_analysis/ correlation_analysis/ .ipynb_checkpoints/ correlation_anaysis_AD-c heckpoint.ipynb - repository limit reached (2,000 files or 30 MB): the rest is at the source (16 files)
Code availability
GitHub: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 39 scripts, each with its path and the digest of its content;
- 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- synapse.org/
synapse:syn3219045 , at Synapse; found in “Data availability” - zenodo:19543885, at Zenodo; found in “Data availability”
Data availability
RNA-Seq data: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 2 keywords, 1 funder, 44 references.
Cite
This paper
Anwer, D., A, A., Marchi, A., Kerkhoven, E., & Polster, A. (2026). Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease. npj aging, 12(1), 95. https://
BibTeX
@article{anwer2026networ
author = {Anwer, Danish and A, Arina and Marchi, Agata and Kerkhoven, Eduard and Polster, Annikka},
title = {{Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease}},
journal = {npj aging},
year = {2026},
month = jul,
volume = {12},
number = {1},
pages = {95},
publisher = {Nature Publishing Group},
issn = {2731-6068},
doi = {10.1038/
url = {https://
pmid = {42463670},
pmcid = {PMC13376175}
}
RIS
TY - JOUR
AU - Anwer, Danish
AU - A, Arina
AU - Marchi, Agata
AU - Kerkhoven, Eduard
AU - Polster, Annikka
TI - Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease
T2 - npj aging
J2 - NPJ Aging
PY - 2026
DA - 2026/
VL - 12
IS - 1
SP - 95
SN - 2731-6068
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease",
"container-title": "npj aging",
"author": [
{
"family": "Anwer",
"given": "Danish"
},
{
"family": "A",
"given": "Arina"
},
{
"family": "Marchi",
"given": "Agata"
},
{
"family": "Kerkhoven",
"given": "Eduard"
},
{
"family": "Polster",
"given": "Annikka"
}
],
"container-title-short":
"volume": "12",
"issue": "1",
"page": "95",
"DOI": "10.1038/
"PMID": "42463670",
"PMCID": "PMC13376175",
"ISSN": "2731-6068",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
16
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: DESeq2, NetworkX, Plotly, 10 other tools, 2 references
- [2] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: DESeq2, NetworkX, Plotly, 10 other tools, 1 reference
- [3] doi:10.1038/s43587-026-01207-x [code]
- A microprotein atlas of the human frontal cortex in Alzheimer's disease.Journal: Nature agingIn common: DESeq2, Plotly, reshape2, 6 other tools, synapse.org/synapse:syn3219045, Alzheimer's / dementia
- [4] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: DESeq2, NetworkX, Plotly, 10 other tools
- [5] doi:10.1038/s44318-026-00818-9 [code]
- FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.Journal: The EMBO journalIn common: DESeq2, Plotly, reshape2, 8 other tools, Alzheimer's / dementia, 1 reference
- [6] doi:10.1016/j.xcrm.2026.102651 [code]
- Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.Journal: Cell reports. MedicineIn common: DESeq2, NetworkX, Plotly, 9 other tools
- [7] doi:10.1371/journal.pcbi.1014327 [code]
- Supervised deep learning with gene functional annotation for cell classification.Journal: PLoS computational biologyIn common: DESeq2, reshape2, ggpubr, 8 other tools, Alzheimer's / dementia, 1 reference
- [8] doi:10.1038/s41467-026-72598-z [code]
- Functional impact of genetic background on variable expressivity in neurodevelopmental disorders.Journal: Nature communicationsIn common: DESeq2, NetworkX, ggpubr, 7 other tools, 2 references
- [9] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: DESeq2, NetworkX, reshape2, 9 other tools
- [10] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: DESeq2, NetworkX, reshape2, 9 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 39 scripts, and 7 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:1c60d5bf54a19be8…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
