OSCR

Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Results › Differential expression analysis ↔ src/code_for_lioness_networks_analysis/Sankey_Plot/plots(box_plot_sankey_plots_network_plots).ipynb, lines 1261–1262 · score 0.90 · TNF alpha, sodium symporter activity, hormone activity, extracellular matrix, amino acid, microvillus
  2. [2] § Results › Network metrics differences in AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/nci_netwoks_features_calculation.ipynb, lines 35–83 · score 0.71 · network properties, network metrics, connected component, largest, regulatory networks, modularity
  3. [3] § Results › Network metrics differences in AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/ad_netwoks_features_calculation.ipynb, lines 58–107 · score 0.71 · network properties, network metrics, connected component, largest, regulatory networks, modularity
  4. [4] § Methods › Pathway and cell type-specific enrichment analysis ↔ src/code_for_lioness_networks_analysis/Sankey_Plot/plots(box_plot_sankey_plots_network_plots).ipynb, lines 389–436 · score 0.67 · GO Cellular Component, GO Molecular Function, KEGG, Hallmark, Human, genes
  5. [5] § Methods › Network analysis of AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/nci_netwoks_features_calculation.ipynb, lines 35–83 · score 0.62 · network metrics, properties, largest, regulatory networks, modularity, nestedness
  6. [6] § Methods › Network analysis of AD-specific and NCI-specific GRNs ↔ src/code_for_lioness_networks_analysis/calculating_graph_features/ad_netwoks_features_calculation.ipynb, lines 58–107 · score 0.62 · network metrics, properties, largest, regulatory networks, modularity, nestedness
  7. [7] § Methods › Data preprocessing ↔ src/code_for_data_preprocessing/combat_seq_batch_removal.ipynb, lines 1–15 · score 0.53 · Combat Seq, batch, preprocessing

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 1,329 lines · 36 KB · no license · 2 matches

  1. # %%
  2. import pandas as pd
  3. import os
  4. import networkx as nx
  5. from tqdm import tqdm
  6. # %%
  7. preprocessed_ad_single_sample_network_dir_path = '../../../data_preprocessed/grn_output/ad_nci_seperate_network/lioness/ad_preprocessed/'
  8. preprocessed_nci_single_sample_network_dir_path = '../../../data_preprocessed/grn_output/ad_nci_seperate_network/lioness/nci_preprocessed/'
  9. # %%
  10. AD_data = pd.read_csv(preprocessed_ad_single_sample_network_dir_path +'lioness.427_120507.58.csv').rename(columns={'2':'lioness.427_120507.58.csv'.split('.')[1]})
  11. for i in tqdm(os.listdir(preprocessed_ad_single_sample_network_dir_path)):
  12. if i.split('.')[1] in AD_data.columns:
  13. continue
  14. if 'csv' in i:
  15. df = pd.read_csv(preprocessed_ad_single_sample_network_dir_path + i).rename(columns={'2':i.split('.')[1]})
  16. AD_data = pd.merge(AD_data,df, on=['0','1'])
  17. # %%
  18. NCI_data = pd.read_csv(preprocessed_nci_single_sample_network_dir_path +'lioness.713_120531.47.csv').rename(columns={'2':'lioness.713_120531.47.csv'.split('.')[1]})
  19. for i in tqdm(os.listdir(preprocessed_nci_single_sample_network_dir_path)):
  20. if i.split('.')[1] in NCI_data.columns:
  21. continue
  22. if 'csv' in i:
  23. df = pd.read_csv(preprocessed_nci_single_sample_network_dir_path + i).rename(columns={'2':i.split('.')[1]})
  24. NCI_data = pd.merge(NCI_data,df, on=['0','1'])
  25. # %%
  26. AD_data.index = AD_data['0'] + '_' + AD_data['1']
  27. AD_data.drop(columns=['0','1'],inplace = True)
  28. AD_data
  29. # %%
  30. NCI_data.index = NCI_data['0'] + '_' + NCI_data['1']
  31. NCI_data.drop(columns=['0','1'],inplace = True)
  32. NCI_data
  33. # %%
  34. common_edge = set(NCI_data.index).intersection(set(AD_data.index))
  35. AD_data = AD_data.loc[list(common_edge),:]
  36. NCI_data = NCI_data.loc[list(common_edge),:]
  37. # %%
  38. AD_data = AD_data.T
  39. AD_data
  40. # %%
  41. NCI_data = NCI_data.T
  42. NCI_data
  43. # %%
  44. combined = pd.concat([AD_data,NCI_data])
  45. # %%
  46. Label = ['AD' for i in range(len(AD_data))] + ['NCI' for i in range(len(NCI_data))]
  47. Label
  48. # %%
  49. import pandas as pd
  50. from sklearn.decomposition import PCA
  51. import matplotlib.pyplot as plt
  52. import seaborn as sns
  53. # Generate sample data (replace this with your own DataFrame)
  54. combined['Label'] = Label
  55. # Assuming your DataFrame contains only numeric values
  56. # If not, you might need to preprocess your data accordingly
  57. # Perform PCA
  58. pca = PCA(n_components=2)
  59. principal_components = pca.fit_transform(combined.drop(columns = 'Label'))
  60. # Create a new DataFrame with the principal components
  61. pc_df = pd.DataFrame(data=principal_components, columns=['PC1', 'PC2'])
  62. # Concatenate the principal components DataFrame with the original DataFrame
  63. final_df = pd.concat([pc_df], axis=1)
  64. final_df['Label'] = Label
  65. final_df.index = combined.index
  66. # Plot the PCA results
  67. plt.figure(figsize=(8, 8))
  68. sns.scatterplot(x='PC1', y='PC2', data=final_df, hue = Label)
  69. plt.title('PCA Plot')
  70. plt.xlabel('Principal Component 1 (PC1)')
  71. plt.ylabel('Principal Component 2 (PC2)')
  72. plt.show()
  73. # %%
  74. remove = final_df[final_df.PC1> 80].index.tolist()
  75. remove
  76. # %%
  77. outlier_corrected_combined = combined[~combined.index.isin(remove)]
  78. # %%
  79. import pandas as pd
  80. from sklearn.decomposition import PCA
  81. import matplotlib.pyplot as plt
  82. import seaborn as sns
  83. # Generate sample data (replace this with your own DataFrame)
  84. # Assuming your DataFrame contains only numeric values
  85. # If not, you might need to preprocess your data accordingly
  86. # Perform PCA
  87. pca = PCA(n_components=2)
  88. principal_components = pca.fit_transform(outlier_corrected_combined .drop(columns = 'Label'))
  89. # Create a new DataFrame with the principal components
  90. pc_df = pd.DataFrame(data=principal_components, columns=['PC1', 'PC2'])
  91. # Concatenate the principal components DataFrame with the original DataFrame
  92. final_df = pd.concat([pc_df], axis=1)
  93. final_df['Label'] = outlier_corrected_combined.Label.tolist()
  94. final_df.index = outlier_corrected_combined.index
  95. Label = outlier_corrected_combined.Label.tolist()
  96. # Plot the PCA results
  97. plt.figure(figsize=(8, 8))
  98. sns.scatterplot(x='PC1', y='PC2', data=final_df, hue = Label )
  99. plt.title('PCA Plot')
  100. plt.xlabel('Principal Component 1 (PC1)')
  101. plt.ylabel('Principal Component 2 (PC2)')
  102. plt.show()
  103. # %%
  104. l = [
  105. 'ENSG00000166823_ENSG00000178199',
  106. 'ENSG00000157557_ENSG00000121903',
  107. 'ENSG00000028277_ENSG00000039139',
  108. 'ENSG00000081059_ENSG00000255561',
  109. 'ENSG00000167981_ENSG00000143222',
  110. 'ENSG00000198482_ENSG00000156219',
  111. 'ENSG00000188786_ENSG00000083799',
  112. 'ENSG00000168874_ENSG00000139197',
  113. 'ENSG00000139800_ENSG00000150175',
  114. 'ENSG00000079432_ENSG00000214425',
  115. 'ENSG00000162992_ENSG00000139197',
  116. 'ENSG00000256294_ENSG00000272576',
  117. 'ENSG00000169635_ENSG00000234841',
  118. 'ENSG00000125850_ENSG00000198783',
  119. 'ENSG00000078900_ENSG00000062096',
  120. 'ENSG00000120837_ENSG00000123179',
  121. 'ENSG00000079432_ENSG00000188234',
  122. 'ENSG00000188786_ENSG00000162384',
  123. 'ENSG00000251369_ENSG00000265018',
  124. 'ENSG00000234284_ENSG00000149089',
  125. 'ENSG00000198429_ENSG00000107262',
  126. 'ENSG00000188785_ENSG00000122641']
  127. # %%
  128. import pickle
  129. # Path to your pickle file
  130. pickle_file_path = '../../../data_preprocessed/pickle_file/ensemble_to_gene_name.pickle'
  131. # Load the pickle file
  132. with open(pickle_file_path, 'rb') as file:
  133. ensemble_to_gene_name = pickle.load(file)
  134. # %%
  135. def return_gene_name(gene):
  136. try:
  137. if ensemble_to_gene_name[gene] == 'Gene name not found':
  138. return gene
  139. return ensemble_to_gene_name[gene]
  140. except:
  141. return gene
  142. # %%
  143. n = []
  144. for i in l:
  145. m = []
  146. for j in i.split('_'):
  147. m.append(return_gene_name(j))
  148. n.append(m[0] + '-' + m[1])
  149. # %%
  150. n
  151. # %%
  152. l.append('Label')
  153. l
  154. # %%
  155. important_gene_dataframe = combined[l]
  156. important_gene_dataframe.head()
  157. # %%
  158. import pandas as pd
  159. import seaborn as sns
  160. import matplotlib.pyplot as plt
  161. from sklearn.preprocessing import MinMaxScaler
  162. # Set the Seaborn theme for better aesthetics
  163. sns.set_theme(style="whitegrid", palette="muted")
  164. # Assuming 'important_gene_dataframe' is your DataFrame and it has a 'Label' column
  165. # Normalize the entire DataFrame (excluding the 'Label' column) to a 0-1 range
  166. scaler = MinMaxScaler()
  167. important_gene_dataframe_normalized = important_gene_dataframe.copy()
  168. # Apply normalization (only to the gene expression columns, not the 'Label' column)
  169. gene_columns = important_gene_dataframe_normalized.drop(columns=['Label']).columns
  170. important_gene_dataframe_normalized[gene_columns] = scaler.fit_transform(important_gene_dataframe_normalized[gene_columns])
  171. # Filter DataFrames for 'AD' and 'NCI'
  172. ad_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'AD']
  173. nci_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'NCI']
  174. # Drop the 'Label' column and only keep the gene expression values
  175. ad_data = ad_data.drop(columns=['Label'])
  176. nci_data = nci_data.drop(columns=['Label'])
  177. # List of genes (columns) to iterate over
  178. genes = ad_data.columns
  179. # Number of genes
  180. num_genes = len(genes)
  181. # Determine the number of rows and columns for subplots (depending on the number of genes)
  182. ncols = 12 # Set the number of columns (adjust based on your preference)
  183. nrows = (num_genes + ncols - 1) // ncols # Automatically adjust the number of rows
  184. # Create a figure and axes for subplots
  185. fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(50, 10 * nrows))
  186. axes = axes.flatten() # Flatten axes array for easy indexing
  187. # Loop through each gene and create a boxplot on a subplot
  188. for idx, gene in enumerate(genes):
  189. ax = axes[idx]
  190. # Prepare the data for the current gene (AD vs NCI)
  191. data = pd.DataFrame({
  192. 'Group': ['AD'] * len(ad_data[gene]) + ['NCI'] * len(nci_data[gene]),
  193. 'Expression': list(ad_data[gene]) + list(nci_data[gene])
  194. })
  195. # Create a boxplot for the current gene
  196. sns.boxplot(data=data, x='Group', y='Expression', palette='pastel', ax=ax)
  197. gene = n[idx]
  198. # Add a title and custom labels
  199. ax.set_title(f'{gene}', fontsize=14, weight='bold')
  200. ax.set_xlabel('Group', fontsize=12)
  201. ax.set_ylabel('Normalized Expression (0-1)', fontsize=12)
  202. # Add gridlines for better readability
  203. ax.grid(True, axis='y', linestyle='--', alpha=0.6)
  204. # Remove any empty subplots if the number of genes is not a perfect multiple of ncols
  205. for idx in range(num_genes, len(axes)):
  206. fig.delaxes(axes[idx])
  207. # Adjust layout for better spacing between subplots
  208. plt.tight_layout()
  209. # Save the figure as an SVG with high resolution
  210. fig.savefig('gene_expression_boxplots.svg', format='svg', dpi=600)
  211. # Show the plot
  212. plt.show()
  213. # %%
  214. l.remove('Label')
  215. plt.figure(figsize=(8, 6)) # Optional: adjust the figure size
  216. sns.heatmap(important_gene_dataframe[l], annot=False, cmap='YlGnBu') # 'annot=True' adds annotations inside the heatmap cells
  217. plt.title('Heatmap of DataFrame')
  218. plt.show()
  219. # %%
  220. AD_panda = pd.read_csv('../../../data_preprocessed/grn_output/ad_nci_seperate_network/panda/output_panda_ad_csv.csv')
  221. # %%
  222. AD_panda.sort_values(by = 'force', ascending=False)
  223. # %%
  224. import pandas as pd
  225. import seaborn as sns
  226. import matplotlib.pyplot as plt
  227. import numpy as np
  228. # Set the plot size
  229. plt.figure(figsize=(8, 6))
  230. # Plot percentile plot using seaborn
  231. sns.ecdfplot(AD_panda['force'], color='skyblue')
  232. percentile_95 = np.percentile(AD_panda['force'], 95)
  233. plt.axvline(x=percentile_95, color='red', linestyle='--', label='90th Percentile')
  234. plt.axhline(y=0.9, color='green', linestyle='--')
  235. # Annotate the point on the plot
  236. plt.annotate(f'({percentile_95:.2f}, 0.95)', xy=(percentile_95, 0.9), xytext=(percentile_95 + 0.1, 0.9),
  237. arrowprops=dict(facecolor='black', arrowstyle='->'))
  238. # Display the plot
  239. plt.title('Percentile Plot (CDF)')
  240. plt.xlabel('Data')
  241. plt.ylabel('Cumulative Probability')
  242. plt.show()
  243. # %%
  244. AD_panda_top_5 = AD_panda[AD_panda.force>=3.71]
  245. AD_panda_top_5
  246. # %%
  247. list_of_tf = list(set(AD_panda_top_5.tf))
  248. AD_panda_top_5_bipartite_complete = AD_panda_top_5[~AD_panda_top_5.gene.isin(list_of_tf)]
  249. # %%
  250. import pandas as pd
  251. import networkx as nx
  252. import matplotlib.pyplot as plt
  253. from tqdm import tqdm
  254. # Create a weighted bipartite graph
  255. AD_BG_bipartite = nx.Graph()
  256. AD_BG_bipartite.add_nodes_from(AD_panda_top_5_bipartite_complete['tf'], bipartite = 0)
  257. AD_BG_bipartite.add_nodes_from(AD_panda_top_5_bipartite_complete['gene'], bipartite = 1)
  258. # Add weighted edges
  259. edges = [(row['tf'], row['gene'], {'weight': row['force']}) for idx, row in tqdm(AD_panda_top_5_bipartite_complete.iterrows())]
  260. AD_BG_bipartite.add_edges_from(edges)
  261. # %%
  262. relevant_genes = l.copy()
  263. # %%
  264. relevant_genes
  265. # %%
  266. l = []
  267. for i in relevant_genes:
  268. l.append(i.split('_')[0])
  269. l.append(i.split('_')[1])
  270. # %%
  271. import gseapy as gp
  272. from gseapy import Msigdb
  273. names = gp.get_library_name(organism='Human')
  274. # %%
  275. l[0:5]
  276. # %%
  277. def return_gene_name(gene):
  278. try:
  279. if ensemble_to_gene_name[gene] == 'Gene name not found':
  280. return gene
  281. return ensemble_to_gene_name[gene]
  282. except:
  283. return gene
  284. # %%
  285. 0%2 != 0
  286. # %%
  287. 2%2 != 0
  288. # %%
  289. df = pd.DataFrame()
  290. data_dict = dict()
  291. for i,k in tqdm(enumerate(l)):
  292. if i%2 != 0:
  293. continue
  294. m = []
  295. try:
  296. #print(k)
  297. #print(l[i+1])
  298. m.append(k)
  299. m.extend(list(AD_BG_bipartite[k]))
  300. m.extend(list(AD_BG_bipartite[l[i+1]]))
  301. except:
  302. print(k)
  303. m = set(m)
  304. print(len(m))
  305. gene_set = [return_gene_name(gene_id) for gene_id in m]
  306. gene_set = list(set(gene_set))
  307. try:
  308. gene_set.remove('Gene name not found')
  309. except:
  310. print('skip')
  311. filtered_gene_set = []
  312. k = return_gene_name(k)
  313. h = return_gene_name(l[i+1])
  314. for p,j in enumerate(gene_set):
  315. if type(j) == type(gene_set[0]):
  316. filtered_gene_set.append(j)
  317. else:
  318. continue
  319. print(len(filtered_gene_set))
  320. enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
  321. df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
  322. df2 = df2[df2['Adjusted P-value']<0.05]
  323. data_dict[f'{k}-{h}'] = list(df2.Term)
  324. print(f'{k}-{h}')
  325. df = pd.concat([df,df2])
  326. # %%
  327. df
  328. # %%
  329. data_dict.keys()
  330. # %%
  331. import pandas as pd
  332. from tabulate import tabulate
  333. from openpyxl import Workbook, load_workbook
  334. from openpyxl.styles import Border, Side, Font, PatternFill
  335. # Input dictionary
  336. data = data_dict
  337. # Extract unique pathways
  338. pathways = sorted(set(p for v in data.values() for p in v))
  339. columns = list(data.keys())
  340. df = pd.DataFrame(index=pathways, columns=columns)
  341. column_widths = {
  342. "A": 80, # Increase column A width
  343. "B": 25, # Increase column B width
  344. "C": 25, # Increase column C width
  345. "D": 25,
  346. "E":25,
  347. "F": 25, # Increase column A width
  348. "G": 25, # Increase column B width
  349. "H": 25, # Increase column C width
  350. "I": 25,
  351. "K":25,
  352. "L": 25, # Increase column A width
  353. "M": 25, # Increase column B width
  354. "N": 25, # Increase column C width
  355. "O": 25,
  356. "P":25,
  357. "Q": 25, # Increase column A width
  358. "R": 25, # Increase column B width
  359. "S": 25, # Increase column C width
  360. "T": 25,
  361. "U":25,
  362. "V": 25, # Increase column A width
  363. "W": 25 #Increase column D width
  364. }
  365. # Function to insert tick mark
  366. def tick():
  367. return "✔"
  368. # Fill DataFrame with tick marks
  369. for key, values in data.items():
  370. for value in values:
  371. df.loc[value, key] = tick()
  372. df.fillna("", inplace=True)
  373. # Define styles
  374. thin_border = Border(left=Side(style='thin'),
  375. right=Side(style='thin'),
  376. top=Side(style='thin'),
  377. bottom=Side(style='thin'))
  378. bold_font = Font(bold=True)
  379. gray_fill = PatternFill(start_color="D3D3D3", end_color="D3D3D3", fill_type="solid")
  380. # Set chunk sizes
  381. columns_per_sheet = 25
  382. rows_per_sheet = 10000
  383. # Split into chunks of 8 columns and 28 rows
  384. total_col_chunks = (len(df.columns) // columns_per_sheet) + (1 if len(df.columns) % columns_per_sheet else 0)
  385. total_row_chunks = (len(df) // rows_per_sheet) + (1 if len(df) % rows_per_sheet else 0)
  386. file_count = 1
  387. for col_chunk in range(total_col_chunks):
  388. for row_chunk in range(total_row_chunks):
  389. start_col = col_chunk * columns_per_sheet
  390. end_col = start_col + columns_per_sheet
  391. start_row = row_chunk * rows_per_sheet
  392. end_row = start_row + rows_per_sheet
  393. df_chunk = df.iloc[start_row:end_row, start_col:end_col]
  394. # Ensure the chunk has the correct dimensions
  395. while len(df_chunk.columns) < columns_per_sheet:
  396. df_chunk[f"Extra_{len(df_chunk.columns)+1}"] = ""
  397. excel_path = f"/12tb_dsk2/danish/GRN_new_21_feb_2025/data_preprocessed/grn_analysis_output/gsea_excel_sheets/pathway_table_part_{file_count}.xlsx"
  398. df_chunk.to_excel(excel_path)
  399. # Load the workbook and apply formatting
  400. wb = load_workbook(excel_path)
  401. ws = wb.active
  402. # Freeze headers
  403. ws.freeze_panes = "B2" # Freeze first row and column
  404. # Apply styles
  405. for row in ws.iter_rows():
  406. for cell in row:
  407. cell.border = thin_border
  408. if cell.value == "✔":
  409. cell.font = bold_font
  410. cell.fill = gray_fill
  411. # Increase row height
  412. for row in ws.iter_rows():
  413. ws.row_dimensions[row[0].row].height = 20
  414. for col_letter, width in column_widths.items():
  415. ws.column_dimensions[col_letter].width = width
  416. wb.save(excel_path)
  417. file_count += 1
  418. print("Excel files created successfully! ✅")
  419. # %%
  420. df
  421. # %%
  422. # Convert DataFrame to HTML
  423. html_path = "temp.html"
  424. df.to_html(html_path, index=False)
  425. # Convert HTML to PDF
  426. pdf_path = "output.pdf"
  427. pdfkit.from_file(html_path, pdf_path)
  428. print(f"Excel file converted to PDF at {pdf_path}")
  429. # %%
  430. # Convert DataFrame to HTML
  431. html_path = "temp.html"
  432. df.to_html(html_path, index=False)
  433. # Convert HTML to PDF
  434. pdf_path = "output.pdf"
  435. pdfkit.from_file(html_path, pdf_path)
  436. print(f"Excel file converted to PDF at {pdf_path}")
  437. # %%
  438. hub_tf= ['ENSG00000186300','ENSG00000256294','ENSG00000234284','ENSG00000188785']
  439. # %%
  440. hub_tf = set(l).intersection(set(hub_tf))
  441. # %%
  442. hub_tf = list(hub_tf)
  443. hub_tf
  444. # %%
  445. relevant_genes
  446. # %%
  447. hub_tf = ['ENSG00000188785', 'ENSG00000234284', 'ENSG00000256294']
  448. # %%
  449. hub_attached_genes = ['ENSG00000122641','ENSG00000149089','ENSG00000272576'] # manually checked
  450. # %%
  451. m = []
  452. for k,i in enumerate(hub_tf):
  453. print(return_gene_name(hub_tf[k]))
  454. print(return_gene_name(hub_attached_genes[k]))
  455. m.append(f'{return_gene_name(hub_tf[k])}-{return_gene_name(hub_attached_genes[k])}')
  456. # %%
  457. m
  458. # %%
  459. m.append('Label')
  460. # %%
  461. n
  462. # %%
  463. n.append('Label')
  464. n
  465. # %%
  466. important_gene_dataframe.columns = n
  467. # %%
  468. important_gene_dataframe
  469. # %%
  470. df = important_gene_dataframe[m].copy()
  471. # Splitting data by labels
  472. df
  473. # %%
  474. important_gene_dataframe.columns
  475. # %%
  476. ad_data = df[df['Label'] == 'AD'].drop(columns='Label')
  477. nci_data = df[df['Label'] == 'NCI'].drop(columns='Label')
  478. from scipy import stats
  479. # Initialize dictionary to store p-values for each column
  480. p_values = {}
  481. # Perform median test for each column
  482. for column in ad_data.columns:
  483. # Get the data for the current column in each group
  484. group1 = ad_data[column]
  485. group2 = nci_data[column]
  486. # Perform the median test
  487. stat, p_value, med, table = stats.median_test(group1, group2)
  488. # Store the p-value for the column
  489. p_values[column] = p_value
  490. # Output the p-values
  491. print(p_values)
  492. # %%
  493. ad_data[['ZNF548-INHBA','ZNF879-APIP']]
  494. # %%
  495. # %%
  496. import pandas as pd
  497. import seaborn as sns
  498. import matplotlib.pyplot as plt
  499. from sklearn.preprocessing import MinMaxScaler
  500. # Set the Seaborn theme for better aesthetics
  501. sns.set_theme(style="whitegrid", palette="muted")
  502. # Assuming 'important_gene_dataframe' is your DataFrame and it has a 'Label' column
  503. # Normalize the entire DataFrame (excluding the 'Label' column) to a 0-1 range
  504. scaler = MinMaxScaler()
  505. important_gene_dataframe_normalized = important_gene_dataframe[m].copy()
  506. # Apply normalization (only to the gene expression columns, not the 'Label' column)
  507. gene_columns = important_gene_dataframe_normalized.drop(columns=['Label']).columns
  508. important_gene_dataframe_normalized[gene_columns] = scaler.fit_transform(important_gene_dataframe_normalized[gene_columns])
  509. # Filter DataFrames for 'AD' and 'NCI'
  510. ad_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'AD']
  511. nci_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'NCI']
  512. # Drop the 'Label' column and only keep the gene expression values
  513. ad_data = ad_data.drop(columns=['Label'])
  514. nci_data = nci_data.drop(columns=['Label'])
  515. # List of genes (columns) to iterate over
  516. genes = ad_data.columns
  517. # Number of genes
  518. num_genes = len(genes)
  519. ad_data = remove_outliers(ad_data)
  520. nci_data = remove_outliers(nci_data)
  521. # Determine the number of rows and columns for subplots (depending on the number of genes)
  522. ncols = 12 # Set the number of columns (adjust based on your preference)
  523. nrows = (num_genes + ncols - 1) // ncols # Automatically adjust the number of rows
  524. # Create a figure and axes for subplots
  525. fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(50, 10 * nrows))
  526. axes = axes.flatten() # Flatten axes array for easy indexing
  527. # Loop through each gene and create a boxplot on a subplot
  528. for idx, gene in enumerate(genes):
  529. ax = axes[idx]
  530. # Prepare the data for the current gene (AD vs NCI)
  531. data = pd.DataFrame({
  532. 'Group': ['AD'] * len(ad_data[gene]) + ['NCI'] * len(nci_data[gene]),
  533. 'Expression': list(ad_data[gene]) + list(nci_data[gene])
  534. })
  535. # Create a boxplot for the current gene
  536. sns.boxplot(data=data, x='Group', y='Expression', palette='pastel', ax=ax)
  537. gene = m[idx]
  538. # Add a title and custom labels
  539. ax.set_title(f'{gene}', fontsize=14, weight='bold')
  540. ax.set_xlabel('Group', fontsize=12)
  541. ax.set_ylabel('Normalized Expression (0-1)', fontsize=12)
  542. # Add gridlines for better readability
  543. max_y = max(data['Expression']) + 0.05
  544. line_y = max_y + 0.02 # Line position
  545. cap_height = 0.01 # Height of vertical caps
  546. # Add horizontal line for significance comparison
  547. ax.hlines(y=line_y, xmin=0, xmax=1, color='black', linewidth=1.5)
  548. # Add vertical caps at both ends of the horizontal line
  549. ax.vlines(x=0, ymin=line_y - cap_height, ymax=line_y, color='black', linewidth=1.5)
  550. ax.vlines(x=1, ymin=line_y - cap_height, ymax=line_y, color='black', linewidth=1.5)
  551. # Display p-value above the plot
  552. ax.text(0.5, line_y + 0.02, f"p = {p_value:.3f}", ha='center', va='bottom', fontsize=12, color='black')
  553. # Gridlines for better readability
  554. ax.grid(True, axis='y', linestyle='--', alpha=0.6)
  555. # Remove any empty subplots if the number of genes is not a perfect multiple of ncols
  556. for idx in range(num_genes, len(axes)):
  557. fig.delaxes(axes[idx])
  558. # Adjust layout for better spacing between subplots
  559. plt.tight_layout()
  560. # Save the figure as an SVG with high resolution
  561. fig.savefig('/12tb_dsk2/danish/GRN_new_21_feb_2025/data_preprocessed/grn_analysis_output/plots/gene_expression_boxplots_subset_after_removing_outliers.svg', format='svg', dpi=300)
  562. # Show the plot
  563. plt.show()
  564. # %%
  565. import pandas as pd
  566. import seaborn as sns
  567. import matplotlib.pyplot as plt
  568. from sklearn.preprocessing import MinMaxScaler
  569. # Set the Seaborn theme for better aesthetics
  570. sns.set_theme(style="whitegrid", palette="muted")
  571. # Assuming 'important_gene_dataframe' is your DataFrame and it has a 'Label' column
  572. scaler = MinMaxScaler()
  573. # Copy the DataFrame and exclude the 'Label' column from normalization
  574. important_gene_dataframe_normalized = important_gene_dataframe.copy()
  575. gene_columns = important_gene_dataframe.drop(columns=['Label']).columns
  576. important_gene_dataframe_normalized[gene_columns] = scaler.fit_transform(important_gene_dataframe[gene_columns])
  577. # Filter DataFrames for 'AD' and 'NCI'
  578. ad_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'AD'].copy()
  579. nci_data = important_gene_dataframe_normalized[important_gene_dataframe_normalized['Label'] == 'NCI'].copy()
  580. # Remove the 'Label' column to work only with gene expression values
  581. ad_data = ad_data.drop(columns=['Label'])
  582. nci_data = nci_data.drop(columns=['Label'])
  583. # Function to remove outliers using IQR
  584. def remove_outliers(df):
  585. Q1 = df.quantile(0.25) # First quartile
  586. Q3 = df.quantile(0.75) # Third quartile
  587. IQR = Q3 - Q1 # Interquartile range
  588. lower_bound = Q1 - 1.5 * IQR
  589. upper_bound = Q3 + 1.5 * IQR
  590. return df[~((df < lower_bound) | (df > upper_bound)).any(axis=1)]
  591. # Apply outlier removal
  592. ad_data_clean = remove_outliers(ad_data)
  593. nci_data_clean = remove_outliers(nci_data)
  594. # List of genes to iterate over
  595. genes = ad_data_clean.columns
  596. num_genes = len(genes)
  597. # Determine subplot layout
  598. ncols = 12
  599. nrows = (num_genes + ncols - 1) // ncols # Auto-adjust rows
  600. # Create figure
  601. fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(50, 10 * nrows))
  602. axes = axes.flatten()
  603. # Loop through each gene to create boxplots
  604. for idx, gene in enumerate(genes):
  605. ax = axes[idx]
  606. # Prepare data for plotting
  607. data = pd.DataFrame({
  608. 'Group': ['AD'] * len(ad_data_clean[gene]) + ['NCI'] * len(nci_data_clean[gene]),
  609. 'Expression': list(ad_data_clean[gene]) + list(nci_data_clean[gene])
  610. })
  611. # Create boxplot
  612. sns.boxplot(data=data, x='Group', y='Expression', palette='pastel', ax=ax)
  613. # Set title and labels
  614. ax.set_title(f'{gene}', fontsize=14, weight='bold')
  615. ax.set_xlabel('Group', fontsize=12)
  616. ax.set_ylabel('Normalized Expression (0-1)', fontsize=12)
  617. # Grid for better readability
  618. ax.grid(True, axis='y', linestyle='--', alpha=0.6)
  619. # Remove empty subplots
  620. for idx in range(num_genes, len(axes)):
  621. fig.delaxes(axes[idx])
  622. # Adjust layout
  623. plt.tight_layout()
  624. # Show plot
  625. plt.show()
  626. # %%
  627. m
  628. # %%
  629. df = pd.DataFrame()
  630. data_dict = dict()
  631. for i,k in tqdm(enumerate(l)):
  632. if i%2 != 0:
  633. continue
  634. m = []
  635. try:
  636. #print(k)
  637. #print(l[i+1])
  638. m.append(k)
  639. m.extend(list(AD_BG_bipartite[k]))
  640. m.extend(list(AD_BG_bipartite[l[i+1]]))
  641. except:
  642. print(k)
  643. m = set(m)
  644. print(len(m))
  645. gene_set = [return_gene_name(gene_id) for gene_id in m]
  646. gene_set = list(set(gene_set))
  647. try:
  648. gene_set.remove('Gene name not found')
  649. except:
  650. print('skip')
  651. filtered_gene_set = []
  652. k = return_gene_name(k)
  653. h = return_gene_name(l[i+1])
  654. for p,j in enumerate(gene_set):
  655. if type(j) == type(gene_set[0]):
  656. filtered_gene_set.append(j)
  657. else:
  658. continue
  659. print(len(filtered_gene_set))
  660. enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
  661. df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
  662. df2 = df2[df2['Adjusted P-value']<0.005]
  663. data_dict[f'{k}-{h}'] = list(df2.Term)
  664. print(f'{k}-{h}')
  665. df = pd.concat([df,df2])
  666. # %%
  667. # 'GO_Biological_Process_2023'
  668. # %%
  669. df.Gene_set.value_counts()
  670. # %%
  671. df
  672. # %%
  673. import plotly.graph_objects as go
  674. # Data preparation as before
  675. sources = []
  676. targets = []
  677. for gene, pathways in data_dict.items():
  678. for pathway in pathways:
  679. sources.append(gene)
  680. targets.append(pathway)
  681. # Create a mapping of genes and pathways to unique indices
  682. unique_items = list(set(sources + targets))
  683. item_indices = {item: idx for idx, item in enumerate(unique_items)}
  684. # Create lists of indices for the Sankey diagram
  685. source_indices = [item_indices[gene] for gene in sources]
  686. target_indices = [item_indices[pathway] for pathway in targets]
  687. # Define the Sankey diagram figure
  688. fig = go.Figure(go.Sankey(
  689. node=dict(
  690. pad=15,
  691. thickness=20,
  692. line=dict(color="black", width=0.5),
  693. label=unique_items,
  694. align='right'
  695. ),
  696. link=dict(
  697. source=source_indices,
  698. target=target_indices,
  699. value=[1] * len(sources), # All links have equal value
  700. )
  701. ))
  702. # Update layout with increased figure size
  703. fig.update_layout(
  704. title_text="Gene to Pathway Sankey Diagram",
  705. font_size=12,
  706. width=900, # Increase width
  707. height=1550, # Increase height
  708. )
  709. # Save the figure as SVG
  710. fig.write_image("/12tb_dsk2/danish/GRN_new_21_feb_2025/data_preprocessed/grn_analysis_output/plots/sankey_enriched_plots.svg", format="svg", scale=1)# Show the figure
  711. fig.show()
  712. # %%
  713. set(df.Term)
  714. # %%
  715. df = pd.DataFrame()
  716. data_dict = dict()
  717. for i,k in tqdm(enumerate(l)):
  718. if i%2 != 0:
  719. continue
  720. m = []
  721. try:
  722. #print(k)
  723. #print(l[i+1])
  724. m.append(k)
  725. m.extend(list(AD_BG_bipartite[k]))
  726. m.extend(list(AD_BG_bipartite[l[i+1]]))
  727. except:
  728. print(k)
  729. m = set(m)
  730. print(len(m))
  731. gene_set = [return_gene_name(gene_id) for gene_id in m]
  732. gene_set = list(set(gene_set))
  733. try:
  734. gene_set.remove('Gene name not found')
  735. except:
  736. print('skip')
  737. filtered_gene_set = []
  738. k = return_gene_name(k)
  739. h = return_gene_name(l[i+1])
  740. for p,j in enumerate(gene_set):
  741. if type(j) == type(gene_set[0]):
  742. filtered_gene_set.append(j)
  743. else:
  744. continue
  745. print(len(filtered_gene_set))
  746. enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
  747. df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
  748. df2 = df2[df2['Adjusted P-value']<0.05]
  749. data_dict[f'{k}-{h}'] = list(df2.Term)
  750. print(f'{k}-{h}')
  751. df = pd.concat([df,df2])
  752. # %%
  753. data_dict.keys()
  754. # %%
  755. data_dict['MTF1-CYLD']
  756. # %%
  757. relevant_genes
  758. # %%
  759. AD_related_genes = pd.read_csv('../../../data_preprocessed/previously_researched_data/gene-list.csv')
  760. AD_related_genes
  761. # %%
  762. gene_name_to_ensemble = {v: k for k, v in ensemble_to_gene_name.items()}
  763. # %%
  764. def return_ensemble_name(gene_name):
  765. try:
  766. return gene_name_to_ensemble[gene_name]
  767. except:
  768. return gene_name
  769. # %%
  770. AD_related_genes_list = [return_ensemble_name(gene_name) for gene_name in AD_related_genes['Gene Symbol'].tolist()]
  771. # %%
  772. AD_related_genes_list[0:10]
  773. # %%
  774. s = set()
  775. for n in range(0,len(l),2):
  776. k = l[n]
  777. connected_nodes = list(AD_BG_bipartite.neighbors(k))
  778. genes_l = set(connected_nodes).intersection(set(AD_related_genes_list))
  779. x = len(set(connected_nodes).intersection(set(AD_related_genes_list)))
  780. if x>0:
  781. print(k)
  782. print(len(connected_nodes))
  783. print(len(genes_l))
  784. print(len(genes_l)/len(connected_nodes))
  785. print(genes_l)
  786. s.update(genes_l)
  787. # %%
  788. len(s)
  789. # %%
  790. len(s)/948
  791. # %%
  792. import networkx as nx
  793. # Assuming l is your list of nodes and AD_BG_bipartite is your bipartite graph
  794. # Initialize a set to store the nodes to be included in the subgraph
  795. nodes_to_include = set(l)
  796. for n in range(0,len(l),2):
  797. k = l[n]
  798. connected_nodes = list(AD_BG_bipartite.neighbors(k))
  799. nodes_to_include.update(connected_nodes)
  800. # Now create the subgraph containing only the nodes in nodes_to_include
  801. subgraph = AD_BG_bipartite.subgraph(nodes_to_include).copy()
  802. # subgraph now contains the nodes in 'l' and their neighbors
  803. # %%
  804. highlighted_edges = []
  805. for i in relevant_genes:
  806. tf = i.split('_')[0]
  807. gene = i.split('_')[1]
  808. highlighted_edges.append((tf, gene))
  809. # %%
  810. import networkx as nx
  811. import matplotlib.pyplot as plt
  812. from adjustText import adjust_text
  813. # Assuming l is your list of nodes and AD_BG_bipartite is your bipartite graph
  814. # Initialize a set to store the nodes to be included in the subgraph
  815. nodes_to_include = set()
  816. s = []
  817. # Add neighbors to nodes_to_include
  818. for n in range(0, len(l), 2):
  819. k = l[n]
  820. nodes_to_include.update(set([k]))
  821. s.append(l[n])
  822. connected_nodes = list(AD_BG_bipartite.neighbors(k))
  823. genes_l = set(connected_nodes).intersection(set(AD_related_genes_list))
  824. genes_l.update(l)
  825. nodes_to_include.update(genes_l)
  826. # Now create the subgraph containing only the nodes in nodes_to_include
  827. subgraph = AD_BG_bipartite.subgraph(nodes_to_include).copy()
  828. # Position nodes using a layout (for example, spring layout)
  829. pos = nx.spring_layout(subgraph)
  830. # Draw nodes and edges
  831. plt.figure(figsize=(20, 20))
  832. node_sizes = []
  833. # Draw the nodes
  834. # Nodes from list 'l' will be squares, others will be circles
  835. node_shapes = []
  836. for node in subgraph.nodes():
  837. if node in s:
  838. node_shapes.append('s')
  839. node_sizes.append(380) # Square for nodes in 'l'
  840. else:
  841. node_shapes.append('o')
  842. node_sizes.append(250) # Circle for connected nodes
  843. # Draw the graph with node shapes based on 'l' membership
  844. nx.draw_networkx_edges(subgraph, pos, width=1.0, alpha=0.7, edge_color='blue')
  845. # Draw the nodes with different shapes (squares for 'l' and circles for others)
  846. for i, node in enumerate(subgraph.nodes()):
  847. shape = node_shapes[i]
  848. size = node_sizes[i]
  849. nx.draw_networkx_nodes(subgraph, pos, nodelist=[node], node_size=size, node_color='blue', node_shape=shape)
  850. nx.draw_networkx_nodes(subgraph, pos, nodelist=s, node_size=size, node_color='red',node_shape=shape)
  851. gene_imp = set(l) - set(s)
  852. gene_imp = list(gene_imp)
  853. nx.draw_networkx_nodes(subgraph, pos, nodelist=gene_imp, node_size=450, node_color='red',node_shape='o')
  854. nx.draw_networkx_edges(subgraph, pos, edgelist=highlighted_edges, width=3, alpha=0.8, edge_color='red')
  855. # Labeling (only labels for nodes in 'l')
  856. labels = {node: return_gene_name(node) for node in l} # Label nodes from 'l' only
  857. texts = []
  858. for node, label in labels.items():
  859. text = plt.text(pos[node][0], pos[node][1], label, fontsize=6, fontweight='bold', color='black', ha='left')
  860. texts.append(text)
  861. # Adjust label positions to avoid overlap
  862. adjust_text(texts, arrowprops=dict(arrowstyle="->", color='black', lw=1))
  863. # Finalize plot
  864. plt.title("Subgraph with Adjusted Labels for 'l' Nodes", fontsize=16)
  865. plt.axis('off') # Remove axis
  866. plt.savefig('../../../data_preprocessed/grn_analysis_output/plots/network.svg', format='svg', dpi=600) # Corrected .savefig -> plt.savefig
  867. plt.show()
  868. # %%
  869. labels
  870. # %% [markdown]
  871. # # AD-targeted Genes-collected: https://agora.adknowledgeportal.org/genes/nominated-targets
  872. # %%
  873. # %%
  874. # %%
  875. # %%
  876. # %%
  877. # %%
  878. # %%
  879. # %%
  880. # %%
  881. # %%
  882. # %%
  883. TNF-alpha signaling via NF-kB, structural components like collagen-containing extracellular matrix (GO:0062023) and microvillus (GO:0005902), as well as molecular activities such as hormone activity (GO:0005179)
  884. # %%
  885. 'collagen-containing extracellular matrix (GO:0062023)' in df.Term.tolist()
  886. # %%
  887. df.Term.tolist()
  888. # %%
  889. df = pd.DataFrame()
  890. data_dict = dict()
  891. for i,k in tqdm(enumerate(l)):
  892. if i%2 != 0:
  893. continue
  894. m = []
  895. try:
  896. #print(k)
  897. #print(l[i+1])
  898. m.append(k)
  899. m.extend(list(AD_BG_bipartite[k]))
  900. m.extend(list(AD_BG_bipartite[l[i+1]]))
  901. except:
  902. print(k)
  903. m = set(m)
  904. print(len(m))
  905. gene_set = [return_gene_name(gene_id) for gene_id in m]
  906. gene_set = list(set(gene_set))
  907. try:
  908. gene_set.remove('Gene name not found')
  909. except:
  910. print('skip')
  911. filtered_gene_set = []
  912. k = return_gene_name(k)
  913. h = return_gene_name(l[i+1])
  914. for p,j in enumerate(gene_set):
  915. if type(j) == type(gene_set[0]):
  916. filtered_gene_set.append(j)
  917. else:
  918. continue
  919. print(len(filtered_gene_set))
  920. enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
  921. df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
  922. df2 = df2[df2['Adjusted P-value']<0.05]
  923. data_dict[f'{k}-{h}'] = list(df2.Term)
  924. print(f'{k}-{h}')
  925. df = pd.concat([df,df2])
  926. # %%
  927. df
  928. # %%
  929. 'collagen-containing extracellular matrix (GO:0062023)' in df.Term.tolist()
  930. # %%
  931. diff_exp_enriched = set(['TNF-alpha Signaling via NF-kB', 'Collagen-Containing Extracellular Matrix (GO:0062023)','Microvillus (GO:0005902)', 'Hormone Activity (GO:0005179)', 'Neutral L-amino Acid:Sodium Symporter Activity (GO:0005295)'])
  932. # %%
  933. diff_exp_enriched - set(df.Term.tolist())
  934. # %%
  935. diff_exp_enriched.
  936. # %%
  937. df = pd.DataFrame()
  938. data_dict = dict()
  939. for i,k in tqdm(enumerate(l)):
  940. if i%2 != 0:
  941. continue
  942. m = []
  943. try:
  944. #print(k)
  945. #print(l[i+1])
  946. m.append(k)
  947. m.extend(list(AD_BG_bipartite[k]))
  948. m.extend(list(AD_BG_bipartite[l[i+1]]))
  949. except:
  950. print(k)
  951. m = set(m)
  952. print(len(m))
  953. gene_set = [return_gene_name(gene_id) for gene_id in m]
  954. gene_set = list(set(gene_set))
  955. try:
  956. gene_set.remove('Gene name not found')
  957. except:
  958. print('skip')
  959. filtered_gene_set = []
  960. k = return_gene_name(k)
  961. h = return_gene_name(l[i+1])
  962. for p,j in enumerate(gene_set):
  963. if type(j) == type(gene_set[0]):
  964. filtered_gene_set.append(j)
  965. else:
  966. continue
  967. print(len(filtered_gene_set))
  968. enr = gp.enrichr(gene_list= filtered_gene_set, gene_sets=['GO_Molecular_Function_2023','MSigDB_Hallmark_2020','KEGG_2021_Human','GO_Cellular_Component_2023'], organism='human')
  969. df2 = enr.results.sort_values(by = 'Adjusted P-value')[enr.results.sort_values(by = 'Adjusted P-value')['Adjusted P-value']<=0.05]
  970. df2 = df2[df2['Adjusted P-value']<0.005]
  971. data_dict[f'{k}-{h}'] = list(df2.Term)
  972. print(f'{k}-{h}')
  973. df = pd.concat([df,df2])
  974. # %%
  975. diff_exp_enriched - set(df.Term.tolist())
  976. # %%
  977. # %%
  978. # %%

plots(box_plot_sankey_plots_network_plots).ipynb at commit 3f3f8d9, no license · at the source

Overview

Authors: Danish Anwer1, Arina A1,2, Agata Marchi1, Eduard Kerkhoven1, Annikka Polster1,2
  1. Division of Systems and Synthetic Biology, Department of Life Sciences, Chalmers University of Technology,Gothenburg, Sweden
  2. Department of Microbiology, Oslo University Hospital and University of Oslo,Oslo, Norway
Journal: npj aging, volume 12, issue 1, article 95
Dates: received 16 June 2025; accepted 7 July 2026; published online 16 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41514-026-00443-0 · PMID 42463670 · PMCID PMC13376175 · OpenAlex W7168694839
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: Alzheimer's / dementia (population)
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning, Connectivity, Graphs
Keywords: Neuroscience, Diseases
Topic: Bioinformatics and Genomic Networks (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Area of Advance Health Engineering at Chalmers University of Technology
Citations: not cited yet (Europe PMC); 48 references in the paper

Abstract

Alzheimer’s disease (AD) is a multifactorial neurodegenerative disorder marked by progressive cognitive decline, yet its transcriptional regulatory architecture remains poorly understood. Here, we model sample-specific gene regulatory networks (GRNs) from dorsolateral prefrontal cortex transcriptomes of 87 individuals with AD and 67 non-cognitively impaired (NCI) controls and use a machine learning classifier to detect consistent disease-specific network features. This sample-specific network approach captures inter-individual variation in transcriptional regulation and revealed 22 key transcription factor-gene regulations that distinguish AD from NCI with 96% weighted accuracy. The key transcription factor-gene interactions were enriched in pathways central to AD pathology, including synaptic signalling, mitochondrial function, proteostasis, and neuroinflammation. Network analysis uncovered significant differences in regulatory connectivity between AD and controls, with ZNF225, ZNF849, and ZNF548 emerging as AD-specific regulatory hubs. Moreover, several key regulatory edges showed significant correlations with longitudinal cognitive decline, supporting their clinical relevance. Our findings highlight pervasive transcriptional dysregulation in AD, emphasizing sample-specific GRN modelling’s value in uncovering regulatory mechanisms.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

Polster-lab/Analysis_GRN_ROSMAP

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 3f3f8d9e1bb54c67b99042add4a411fbae4b0127, 3 April 2025
Languages: Jupyter (55)
Size: 62 files, 55 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: 27 notebooks
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: pandas (32 files), NumPy (24 files), Matplotlib (12 files), seaborn (12 files), SciPy (10 files), NetworkX (7 files), Plotly (6 files), scikit-learn (6 files), DESeq2 (1 file), ggplot2 (1 file), ggpubr (1 file), reshape2 (1 file), tidyverse (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
39 files

Code availability

GitHub: https://github.com/Polster-lab/Analysis_GRN_ROSMAP,.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 39 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

RNA-Seq data: https://www.synapse.org/Synapse:syn3219045 The datasets generated during this study are available in the Zenodo repository: 10.5281/zenodo.19543885.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 2 keywords, 1 funder, 44 references.

Cite

This paper

Anwer, D., A, A., Marchi, A., Kerkhoven, E., & Polster, A. (2026). Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease. npj aging, 12(1), 95. https://doi.org/10.1038/s41514-026-00443-0

BibTeX

@article{anwer2026network,
author = {Anwer, Danish and A, Arina and Marchi, Agata and Kerkhoven, Eduard and Polster, Annikka},
title = {{Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease}},
journal = {npj aging},
year = {2026},
month = jul,
volume = {12},
number = {1},
pages = {95},
publisher = {Nature Publishing Group},
issn = {2731-6068},
doi = {10.1038/s41514-026-00443-0},
url = {https://doi.org/10.1038/s41514-026-00443-0},
pmid = {42463670},
pmcid = {PMC13376175}
}

RIS

TY - JOUR
AU - Anwer, Danish
AU - A, Arina
AU - Marchi, Agata
AU - Kerkhoven, Eduard
AU - Polster, Annikka
TI - Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease
T2 - npj aging
J2 - NPJ Aging
PY - 2026
DA - 2026/07/16
VL - 12
IS - 1
SP - 95
SN - 2731-6068
PB - Nature Publishing Group
DO - 10.1038/s41514-026-00443-0
UR - https://doi.org/10.1038/s41514-026-00443-0
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41514-026-00443-0",
"type": "article-journal",
"title": "Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease",
"container-title": "npj aging",
"author": [
{
"family": "Anwer",
"given": "Danish"
},
{
"family": "A",
"given": "Arina"
},
{
"family": "Marchi",
"given": "Agata"
},
{
"family": "Kerkhoven",
"given": "Eduard"
},
{
"family": "Polster",
"given": "Annikka"
}
],
"container-title-short": "NPJ Aging",
"volume": "12",
"issue": "1",
"page": "95",
"DOI": "10.1038/s41514-026-00443-0",
"PMID": "42463670",
"PMCID": "PMC13376175",
"ISSN": "2731-6068",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41514-026-00443-0",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
16
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: DESeq2, NetworkX, Plotly, 10 other tools, 2 references
[2] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: DESeq2, NetworkX, Plotly, 10 other tools, 1 reference
[3] doi:10.1038/s43587-026-01207-x [code]
A microprotein atlas of the human frontal cortex in Alzheimer's disease.
Journal: Nature aging
In common: DESeq2, Plotly, reshape2, 6 other tools, synapse.org/synapse:syn3219045, Alzheimer's / dementia
[4] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: DESeq2, NetworkX, Plotly, 10 other tools
[5] doi:10.1038/s44318-026-00818-9 [code]
FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.
Journal: The EMBO journal
In common: DESeq2, Plotly, reshape2, 8 other tools, Alzheimer's / dementia, 1 reference
[6] doi:10.1016/j.xcrm.2026.102651 [code]
Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.
Journal: Cell reports. Medicine
In common: DESeq2, NetworkX, Plotly, 9 other tools
[7] doi:10.1371/journal.pcbi.1014327 [code]
Supervised deep learning with gene functional annotation for cell classification.
Journal: PLoS computational biology
In common: DESeq2, reshape2, ggpubr, 8 other tools, Alzheimer's / dementia, 1 reference
[8] doi:10.1038/s41467-026-72598-z [code]
Functional impact of genetic background on variable expressivity in neurodevelopmental disorders.
Journal: Nature communications
In common: DESeq2, NetworkX, ggpubr, 7 other tools, 2 references
[9] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: DESeq2, NetworkX, reshape2, 9 other tools
[10] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: DESeq2, NetworkX, reshape2, 9 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.