MLMarker: a machine learning framework for tissue inference and biomarker discovery.
The 34 matches
- [1] § Reuse of pan-cancer dataset (MSV000095036) ↔ Cancer_datasets_Kuster/baseline_comparison_clean.ipynb, lines 79–138 · score 0.85 · parotid gland, small intestine, pituitary gland, colon, duodenum, rectum
- [2] § Baseline atlas comparisons and alternative classifiers › Interpretation analyses ↔ custom_functions.py, lines 248–305 · score 0.84 · cellular components, Human Protein Atlas, molecular functions, biological processes, KEGG, pathways
- [3] § Results › Reprocessed data from cerebral melanoma (PXD007592) ↔ pages/1_QC.py, lines 187–249 · score 0.84 · intensity variation, signal proportion, global sample, intensity distribution, protein intensity, feature space
- [4] § Reuse of pan-cancer dataset (MSV000095036) ↔ Home.py, lines 243–243 · score 0.77 · parotid gland, small intestine, pituitary gland, colon, duodenum, rectum
- [5] § Baseline atlas comparisons and alternative classifiers › Reference atlases ↔ Cancer_datasets_Kuster/baseline_comparison_clean.ipynb, lines 54–77 · score 0.76 · excluding cancer, normalised intensities, UniProt, tissue matrix, training atlas, HPA
- [6] § Results › Circularity and double-dipping ↔ PXD007592_responders_melanoma/03_shap_vs_dea_correlation.ipynb, lines 630–680 · score 0.74 · circular reasoning, reflects cohort, capture tissue, prone, complementary, leakage
- [7] § Results › Penalty factor › Performance across sample types: dense vs. sparse ↔ Missingnes_testing/02_penalty_factor_comparison.ipynb, lines 550–581 · score 0.72 · high dropout, low dropout, sparse biofluid, Penalty benefit, dense tissue, beneficial
- [8] § Penalty factor evaluation by simulated protein missingness › Penalty strategies ↔ Missingnes_testing/penalty_factor_analysis.py, lines 128–144 · score 0.71 · piecewise linear interpolation, Adaptive penalties, low coverage
- [9] § Reprocessed data from cerebrospinal fluid (PXD008029) ↔ Home.py, lines 243–243 · score 0.70 · bone marrow, adipose tissue, pituitary gland, testis, lung, monocytes
- [10] § Reprocessed data from cerebrospinal fluid (PXD008029) ↔ test.ipynb, lines 140–141 · score 0.70 · bone marrow, adipose tissue, pituitary gland, testis, lung, monocytes
- [11] § Results › Penalty factor › Performance across sample types: dense vs. sparse ↔ Missingnes_testing/02_penalty_factor_comparison.ipynb, lines 405–548 · score 0.69 · targeted missingness, sparse samples, dense tissue, loss, missing proteins, MNAR
- [12] § Materials and methods › The MLMarker package › Penalty factor for missing feature effects ↔ old_modules/mlmarker.py, lines 116–165 · score 0.64 · identifies absent proteins, remain unchanged, penalty factor, SHAP, zero, MLMarker
- [13] § Results › Reprocessed data from cerebral melanoma (PXD007592) ↔ PXD007592_responders_melanoma/classify_04_protein_analysis.ipynb, lines 41–127 · score 0.64 · Mann Whitney, good responders, higher brain, poor responders, brain prediction, scores
- [14] § Results › Reprocessed data from cerebral melanoma (PXD007592) ↔ PXD007592_responders_melanoma/05_clusteringgroups.ipynb, lines 224–297 · score 0.63 · response status, Mann Whitney, higher brain, Poor responders, UMAP, clustering
- [15] § Results › Penalty factor ↔ Missingnes_testing/02_penalty_factor_comparison.ipynb, lines 405–548 · score 0.62 · Targeted missingness, detection limits, loss, PXD009021, lost, MNAR
- [16] § Materials and methods › Reprocessing pipeline › Differential expression analysis ↔ PXD007592_responders_melanoma/msqrob_analysis.R, lines 106–138 · score 0.62 · msqrob2, Peptide intensities, aggregated, median, responders, models
- [17] § Materials and methods › The MLMarker package › Annotation utilities and visualisation ↔ custom_functions.py, lines 248–305 · score 0.62 · Human Protein Atlas, tissue expression, Profiler, enrichment, SHAP, score
- [18] § Results › Reprocessed data from cerebral melanoma (PXD007592) ↔ pages/1_QC.py, lines 20–28 · score 0.60 · predefined proteins, missing proteins, artifacts, quality, technical, intensity
- [19] § Baseline atlas comparisons and alternative classifiers › Alternative classification methods ↔ Cancer_datasets_Kuster/baseline_comparison_clean.ipynb, lines 368–415 · score 0.60 · nearest neighbours classification, sample profile, distance, correlation, training, tissue
- [20] § Reuse of pan-cancer dataset (MSV000095036) ↔ Cancer_datasets_Kuster/baseline_comparison_clean.ipynb, lines 79–138 · score 0.60 · healthy tissues, glioma, LN, OE, DLBCL, OSCC
- [21] § Results › Reprocessed data from cerebral melanoma (PXD007592) ↔ PXD007592_responders_melanoma/05_clusteringgroups.ipynb, lines 224–297 · score 0.59 · Mann Whitney, higher brain, poor responders, good responders, brain prediction, UMAP
- [22] § Baseline atlas comparisons and alternative classifiers › Interpretation analyses ↔ pages/2_Visualisations.py, lines 22–31 · score 0.58 · SHapley, exPlanations, Additive, predictions, tissue, proteins
- [23] § Results › Circularity and double-dipping ↔ PXD007592_responders_melanoma/03_shap_vs_dea_correlation.ipynb, lines 87–128 · score 0.57 · MLMarker feature space, proteins outside, log10 transformed, fold change, dipping, Circularity
- [24] § Reprocessed data from cerebrospinal fluid (PXD008029) ↔ PXD008029_CSF/classify_csf.ipynb, lines 211–256 · score 0.56 · adipose tissue, CSF samples, FDR, monocytes, plasma, PXD008029
- [25] § Materials and methods › The MLMarker package › Penalty factor for missing feature effects ↔ mlmarker/explainability.py, lines 146–208 · score 0.56 · optional penalty factor, Absent proteins, SHAP, zero, MLMarker, tissue
- [26] § Results › Circularity and double-dipping ↔ PXD007592_responders_melanoma/03_shap_vs_dea_correlation.ipynb, lines 630–680 · score 0.56 · indicates partial overlap, limited correlation, circular, SHAP, Spearman, MLMarker
- [27] § Results › Penalty factor › Performance across sample types: dense vs. sparse ↔ Missingnes_testing/rank_analysis_heatmap.py, lines 73–136 · score 0.56 · sparse biofluid, dense tissue, dropout rates, heatmap, adaptive, rank
- [28] § Penalty factor evaluation by simulated protein missingness ↔ Missingnes_testing/02_penalty_factor_comparison.ipynb, lines 14–46 · score 0.55 · pituitary gland, target tissue, PXD009021, Simulations, PXD008029, missingness
- [29] § Penalty factor evaluation by simulated protein missingness ↔ Missingnes_testing/penalty_factor_analysis.py, lines 193–205 · score 0.54 · pituitary gland, target tissue, NSAF, ionbot, PXD008029, missingness
- [30] § Results › Penalty factor › Performance across sample types: dense vs. sparse ↔ Missingnes_testing/rank_analysis_heatmap.py, lines 73–136 · score 0.54 · lower ranks, sparse biofluids, dense tissue, missingness, penalty
- [31] § Results › Circularity and double-dipping ↔ PXD007592_responders_melanoma/03_shap_vs_dea_correlation.ipynb, lines 87–128 · score 0.53 · MLMarker feature space, Double dipping, outside, circularity, external, protein
- [32] § Reprocessed data from cerebrospinal fluid (PXD008029) › SHAP analysis of CSF samples within the Streamlit application ↔ PXD008029_CSF/classify_csf.ipynb, lines 291–307 · score 0.52 · bone marrow, pituitary gland, CSF, brain, predicted
- [33] § Results › Differential expression analysis with MSqRob ↔ PXD007592_responders_melanoma/msqrob_analysis.R, lines 141–220 · score 0.51 · log fold change, log10 adjusted, volcano, MSqRob, responder, brain
- [34] § Results › Penalty factor › Performance across sample types: dense vs. sparse ↔ Missingnes_testing/penalty_factor_analysis.py, lines 322–359 · score 0.50 · adaptive piecewise, dropout rates, penalty factor, MNAR, simulating, missingness
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 2,059 lines · 86 KB · no license · 4 matches
- # %%
- # === 1. Setup and Imports ===
- import os
- import warnings
- warnings.filterwarnings('ignore')
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import seaborn as sns
- from scipy.stats import spearmanr
- from tqdm import tqdm
- # Plot settings
- sns.set_style("whitegrid")
- plt.rcParams['figure.figsize'] = (12, 8)
- plt.rcParams['font.size'] = 12
- plt.rcParams['figure.dpi'] = 100
- print("Libraries loaded successfully.")
- # %% [markdown]
- # ## 2. Load Data
- #
- # Load protein expression data, MLMarker predictions, and reference atlases.
- # %%
- # Load protein expression data
- protein_data_path = 'search__Picked_Group_FDR_Outputs__Picked_Group_FDR_Outputs_-_FP_Booster__combined_protein.tsv'
- local_path = 'search__Picked_Group_FDR_Outputs__Picked_Group_FDR_Outputs_-_FP_Booster__combined_protein.tsv'
- protein_df = pd.read_csv(local_path if os.path.exists(local_path) else protein_data_path, sep='\t')
- # Load MLMarker predictions
- mlmarker_predictions = pd.read_csv('Entire_picked_FDR_MLMarker_predictions_dev_penalty1.csv', index_col=0)
- # Load metadata
- metadata = pd.read_csv('MassIVE_Mapping.txt', sep='\t')
- # Load reference atlases
- hpa = pd.read_csv('antibody_atlas.csv')
- pdb = pd.read_csv('proteomicsdb.csv')
- # Load MLMarker training data
- ml_train = pd.read_csv('training_atlas_7592%_10exp.csv.gz', index_col=0)
- print(f"Protein data: {protein_df.shape}")
- print(f"MLMarker predictions: {mlmarker_predictions.shape}")
- print(f"Metadata entries: {len(metadata)}")
- print(f"HPA: {hpa['Uniprot_id'].nunique()} proteins, {hpa['Tissue'].nunique()} tissues")
- print(f"PDB: {pdb['UniProt'].nunique()} proteins, {pdb['tissue'].nunique()} tissues")
- print(f"MLMarker training: {ml_train.shape[0]} samples, {ml_train['tissue_name'].nunique()} tissues")
- # %% [markdown]
- # ## 3. Create Atlas Matrices
- #
- # Convert all data sources to protein × tissue matrices.
- # %%
- # === HPA Atlas Matrix ===
- hpa_matrix = hpa.pivot_table(
- index='Uniprot_id', columns='Tissue', values='Level', aggfunc='mean'
- ).fillna(0)
- print(f"HPA matrix: {hpa_matrix.shape} (proteins × tissues)")
- # === PDB Atlas Matrix (exclude cancer/fluids) ===
- pdb_matrix = pdb.pivot_table(
- index='UniProt', columns='tissue', values='normalised intensity', aggfunc='mean'
- ).fillna(0)
- print(f"PDB matrix: {pdb_matrix.shape} (proteins × tissues)")
- # === MLMarker Training Atlas ===
- metadata_cols = ['tissue_name', 'disease_status', 'fluid', 'cell_type']
- protein_cols = [c for c in ml_train.columns if c not in metadata_cols]
- ml_atlas = ml_train.groupby('tissue_name')[protein_cols].mean().T
- print(f"MLMarker training atlas: {ml_atlas.shape} (proteins × tissues)")
- # %% [markdown]
- # ## 4. Define Tissue Mappings
- #
- # Map cancer cohorts to expected tissue of origin for each atlas.
- # %%
- # Cohorts to evaluate (excluding DLBCL+ as it's redundant with DLBCL)
- eval_cohorts = ['CRC', 'DLBCL', 'Glioma', 'OSCC', 'healthy LN', 'healthy OE']
- # === MLMarker Mappings ===
- cancer_to_mlmarker = {
- 'CRC': 'Colon', 'DLBCL': 'B-cells',
- 'Glioma': 'Brain', 'OSCC': 'Parotid gland',
- 'healthy LN': 'Lymph node', 'healthy OE': 'Parotid gland'
- }
- cancer_to_mlmarker_extended = {
- 'CRC': ['Colon', 'Small intestine', 'Rectum', 'Duodenum', 'Appendix'],
- 'DLBCL': ['B-cells', 'Lymph node', 'Spleen'],
- 'Glioma': ['Brain', 'Pituitary gland'],
- 'OSCC': ['Parotid gland', 'Tonsil'],
- 'healthy LN': ['Lymph node', 'B-cells', 'Spleen'],
- 'healthy OE': ['Parotid gland', 'Tonsil', 'Esophagus']
- }
- # === HPA Mappings ===
- cancer_to_hpa = {
- 'CRC': 'Colon', 'DLBCL': 'Lymph node',
- 'Glioma': 'Cerebral cortex', 'OSCC': 'Salivary gland',
- 'healthy LN': 'Lymph node', 'healthy OE': 'Oral mucosa'
- }
- cancer_to_hpa_extended = {
- 'CRC': ['Colon', 'Small intestine', 'Rectum', 'Duodenum', 'Appendix'],
- 'DLBCL': ['Lymph node', 'Spleen', 'Tonsil', 'Bone marrow'],
- 'Glioma': ['Cerebral cortex', 'Cerebellum', 'Hippocampus', 'Caudate', 'Pituitary gland'],
- 'OSCC': ['Salivary gland', 'Oral mucosa', 'Nasopharynx'],
- 'healthy LN': ['Lymph node', 'Spleen', 'Tonsil'],
- 'healthy OE': ['Oral mucosa', 'Salivary gland', 'Esophagus']
- }
- # === PDB Mappings ===
- cancer_to_pdb = {
- 'CRC': 'Colon', 'DLBCL': 'B-lymphocyte',
- 'Glioma': 'Brain', 'OSCC': 'Salivary gland',
- 'healthy LN': 'Lymph node', 'healthy OE': 'Mouth'
- }
- cancer_to_pdb_extended = {
- 'CRC': ['Colon', 'Colon muscle', 'Gut', 'Small intestine'],
- 'DLBCL': ['B-lymphocyte', 'Lymph node', 'Spleen'],
- 'Glioma': ['Brain', 'Cerebral cortex'],
- 'OSCC': ['Salivary gland', 'Mouth'],
- 'healthy LN': ['Lymph node', 'B-lymphocyte', 'Spleen'],
- 'healthy OE': ['Mouth', 'Salivary gland', 'Esophagus']
- }
- print(f"Cohorts to evaluate: {eval_cohorts}")
- print(f"\nExcluded: Melanoma, HELA, PDAC (no direct healthy tissue equivalent)")
- print(f"Excluded: DLBCL+ (redundant with DLBCL)")
- # %% [markdown]
- # ## 5. Prepare Sample Data
- # %%
- # Get expression columns
- expression_columns = [col for col in protein_df.columns if "MaxLFQ Intensity" in col]
- sample_names = [col.split(' MaxLFQ Intensity')[0] for col in expression_columns]
- # Create sample metadata
- sample_metadata = pd.DataFrame({'sample': sample_names, 'expression_col': expression_columns})
- sample_metadata = sample_metadata.merge(
- metadata[['Cohort', 'Folder name FP_DDA']],
- left_on='sample', right_on='Folder name FP_DDA', how='inner'
- ).drop_duplicates()
- sample_metadata = sample_metadata[sample_metadata['Cohort'].isin(eval_cohorts)]
- # Create expression matrix (log2 transformed)
- valid_cols = sample_metadata['expression_col'].tolist()
- sample_expr = protein_df[['Protein ID'] + valid_cols].set_index('Protein ID')
- sample_expr_log = np.log2(sample_expr + 1)
- # Create MLMarker column mapping
- sample_mapping_mlmarker = {}
- for sample in sample_metadata['sample']:
- matches = [c for c in mlmarker_predictions.columns if sample in c]
- if matches:
- sample_mapping_mlmarker[sample] = matches[0]
- print(f"Samples per cohort:")
- print(sample_metadata['Cohort'].value_counts())
- print(f"\nTotal samples: {len(sample_metadata)}")
- print(f"MLMarker mapping found: {len(sample_mapping_mlmarker)} samples")
- # %% [markdown]
- # ## 6. Define Classification Functions
- # %%
- def atlas_correlation_classify(sample_profile, atlas_matrix):
- """
- Classify a sample using Spearman correlation with atlas tissues.
- Returns: (predicted_tissue, score, all_scores_dict)
- """
- common_proteins = sample_profile.index.intersection(atlas_matrix.index)
- if len(common_proteins) < 10:
- return None, 0, {}
- sample_aligned = sample_profile.loc[common_proteins].values
- atlas_aligned = atlas_matrix.loc[common_proteins]
- correlations = {}
- for tissue in atlas_aligned.columns:
- tissue_profile = atlas_aligned[tissue].values
- if tissue_profile.std() == 0 or sample_aligned.std() == 0:
- correlations[tissue] = 0
- else:
- corr, _ = spearmanr(sample_aligned, tissue_profile)
- correlations[tissue] = corr if not np.isnan(corr) else 0
- if correlations:
- predicted = max(correlations, key=correlations.get)
- return predicted, correlations[predicted], correlations
- return None, 0, {}
- def get_topk_predictions(scores_dict, k=5):
- """Get top-k tissue predictions from scores dictionary."""
- if not scores_dict:
- return []
- return sorted(scores_dict.keys(), key=lambda x: scores_dict[x], reverse=True)[:k]
- print("Classification functions defined.")
- # %% [markdown]
- # ## 7. Run Classification (All 4 Methods)
- # %%
- # Initialize results storage
- results = []
- for _, row in tqdm(sample_metadata.iterrows(), total=len(sample_metadata), desc="Classifying samples"):
- sample_name = row['sample']
- cohort = row['Cohort']
- expr_col = row['expression_col']
- # Get sample profile (non-zero proteins only)
- sample_profile = sample_expr_log[expr_col]
- sample_profile = sample_profile[sample_profile > 0]
- result = {'sample': sample_name, 'cohort': cohort}
- # === 1. MLMarker Model ===
- if sample_name in sample_mapping_mlmarker:
- ml_col = sample_mapping_mlmarker[sample_name]
- ml_scores = mlmarker_predictions[ml_col].to_dict()
- result['mlmarker_pred'] = mlmarker_predictions[ml_col].idxmax()
- else:
- ml_scores = {}
- result['mlmarker_pred'] = None
- true_strict = cancer_to_mlmarker.get(cohort)
- true_extended = cancer_to_mlmarker_extended.get(cohort, [])
- for k in [1, 2, 3, 4, 5]:
- top_k = get_topk_predictions(ml_scores, k)
- result[f'mlmarker_top{k}_strict'] = true_strict in top_k if top_k else False
- result[f'mlmarker_top{k}_extended'] = any(t in top_k for t in true_extended) if top_k else False
- # === 2. MLMarker Training Atlas ===
- pred, _, scores = atlas_correlation_classify(sample_profile, ml_atlas)
- result['ml_training_pred'] = pred
- for k in [1, 2, 3, 4, 5]:
- top_k = get_topk_predictions(scores, k)
- result[f'ml_training_top{k}_strict'] = true_strict in top_k if top_k else False
- result[f'ml_training_top{k}_extended'] = any(t in top_k for t in true_extended) if top_k else False
- # === 3. HPA Atlas ===
- true_hpa_strict = cancer_to_hpa.get(cohort)
- true_hpa_extended = cancer_to_hpa_extended.get(cohort, [])
- pred, _, scores = atlas_correlation_classify(sample_profile, hpa_matrix)
- result['hpa_pred'] = pred
- for k in [1, 2, 3, 4, 5]:
- top_k = get_topk_predictions(scores, k)
- result[f'hpa_top{k}_strict'] = true_hpa_strict in top_k if top_k else False
- result[f'hpa_top{k}_extended'] = any(t in top_k for t in true_hpa_extended) if top_k else False
- # === 4. PDB Atlas ===
- true_pdb_strict = cancer_to_pdb.get(cohort)
- true_pdb_extended = cancer_to_pdb_extended.get(cohort, [])
- pred, _, scores = atlas_correlation_classify(sample_profile, pdb_matrix)
- result['pdb_pred'] = pred
- for k in [1, 2, 3, 4, 5]:
- top_k = get_topk_predictions(scores, k)
- result[f'pdb_top{k}_strict'] = true_pdb_strict in top_k if top_k else False
- result[f'pdb_top{k}_extended'] = any(t in top_k for t in true_pdb_extended) if top_k else False
- results.append(result)
- results_df = pd.DataFrame(results)
- print(f"\nClassification complete! Results: {results_df.shape}")
- # %% [markdown]
- # ## 8. Calculate Top-k Accuracy Summary
- # %%
- # Calculate overall accuracy for each method
- methods = ['mlmarker', 'ml_training', 'hpa', 'pdb']
- method_labels = ['MLMarker', 'MLMarker Training', 'HPA Atlas', 'PDB Atlas']
- summary_data = []
- for method, label in zip(methods, method_labels):
- for k in [1, 2, 3, 4, 5]:
- strict_acc = results_df[f'{method}_top{k}_strict'].mean() * 100
- extended_acc = results_df[f'{method}_top{k}_extended'].mean() * 100
- summary_data.append({
- 'Method': label, 'k': k,
- 'Strict (%)': strict_acc, 'Extended (%)': extended_acc
- })
- summary_df = pd.DataFrame(summary_data)
- # Display Top-1 and Top-5 summary
- print("=" * 70)
- print("OVERALL ACCURACY SUMMARY")
- print("=" * 70)
- print(f"\n{'Method':<20} {'Top-1 Strict':>12} {'Top-1 Extended':>15} {'Top-5 Extended':>15}")
- print("-" * 70)
- for method, label in zip(methods, method_labels):
- t1_strict = results_df[f'{method}_top1_strict'].mean() * 100
- t1_ext = results_df[f'{method}_top1_extended'].mean() * 100
- t5_ext = results_df[f'{method}_top5_extended'].mean() * 100
- print(f"{label:<20} {t1_strict:>11.1f}% {t1_ext:>14.1f}% {t5_ext:>14.1f}%")
- print(f"\nTotal samples: {len(results_df)}")
- # %% [markdown]
- # ## 9. Per-Cohort Accuracy Analysis
- # %%
- # Calculate per-cohort accuracy
- cohort_data = []
- for cohort in eval_cohorts:
- cohort_mask = results_df['cohort'] == cohort
- n_samples = cohort_mask.sum()
- for method, label in zip(methods, method_labels):
- for k in [1, 3, 5]:
- strict_acc = results_df.loc[cohort_mask, f'{method}_top{k}_strict'].mean() * 100
- extended_acc = results_df.loc[cohort_mask, f'{method}_top{k}_extended'].mean() * 100
- cohort_data.append({
- 'Cohort': cohort, 'Method': label, 'k': k, 'n': n_samples,
- 'Strict (%)': strict_acc, 'Extended (%)': extended_acc
- })
- cohort_df = pd.DataFrame(cohort_data)
- # Display per-cohort Top-1 Extended accuracy
- print("\nPer-Cohort Top-1 Extended Accuracy:")
- print("=" * 80)
- pivot = cohort_df[cohort_df['k'] == 1].pivot(index='Cohort', columns='Method', values='Extended (%)')
- pivot = pivot[method_labels] # Reorder columns
- print(pivot.round(1).to_string())
- # %% [markdown]
- # ## 10. K-Nearest Neighbors (KNN) Classification
- #
- # Instead of correlating with tissue centroids (means), KNN classifies based on the k nearest training samples. This preserves sample-level variability and may better capture tissue-specific patterns.
- # %%
- # Prepare MLMarker training data for KNN
- # ml_train has samples as rows, proteins as columns, with tissue_name column
- # Get protein columns (exclude metadata)
- ml_protein_cols = [c for c in ml_train.columns if c not in ['tissue_name', 'disease_status', 'fluid', 'cell_type']]
- ml_train_X = ml_train[ml_protein_cols]
- ml_train_y = ml_train['tissue_name']
- print(f"MLMarker Training data for KNN:")
- print(f" Samples: {len(ml_train_X)}")
- print(f" Proteins: {len(ml_protein_cols)}")
- print(f" Tissues: {ml_train_y.nunique()}")
- print(f" Samples per tissue (top 10):")
- print(ml_train_y.value_counts().head(10))
- # %%
- # KNN Classification Function
- from sklearn.neighbors import KNeighborsClassifier
- from collections import Counter
- def knn_classify_with_topk(sample_profile, train_X, train_y, k=5):
- """
- Classify sample using KNN and return top-k tissue predictions.
- Returns: (predicted_tissue, top_k_tissues, vote_counts)
- """
- # Find common proteins
- common_proteins = sample_profile.index.intersection(train_X.columns)
- if len(common_proteins) < 10:
- return None, [], {}
- # Align data
- sample_aligned = sample_profile.loc[common_proteins].values.reshape(1, -1)
- train_aligned = train_X[common_proteins].values
- train_labels = train_y.values
- # Handle NaN
- sample_aligned = np.nan_to_num(sample_aligned, 0)
- train_aligned = np.nan_to_num(train_aligned, 0)
- # Fit KNN with k neighbors
- knn = KNeighborsClassifier(n_neighbors=min(k, len(train_aligned)), metric='correlation')
- knn.fit(train_aligned, train_labels)
- # Get k nearest neighbors
- distances, indices = knn.kneighbors(sample_aligned)
- neighbor_labels = train_labels[indices[0]]
- # Count votes
- vote_counts = Counter(neighbor_labels)
- # Get top-k unique tissues by vote count
- sorted_tissues = [t for t, _ in vote_counts.most_common()]
- # Get probabilities for all classes
- proba = knn.predict_proba(sample_aligned)[0]
- class_proba = dict(zip(knn.classes_, proba))
- predicted = knn.predict(sample_aligned)[0]
- return predicted, sorted_tissues, class_proba
- print("KNN classification function defined.")
- # %%
- # Run KNN on MLMarker Training Data (sample-level)
- # Test different k values
- k_values = [1, 3, 5, 7, 11]
- knn_results = {f'knn_ml_k{k}_pred': [] for k in k_values}
- knn_results['sample'] = []
- knn_results['cohort'] = []
- # Add top-k accuracy tracking for BOTH strict and extended
- for k in k_values:
- for topk in [1, 2, 3, 4, 5]:
- knn_results[f'knn_ml_k{k}_top{topk}_strict'] = []
- knn_results[f'knn_ml_k{k}_top{topk}_extended'] = []
- print(f"Running KNN classification with k={k_values}...")
- for _, row in tqdm(sample_metadata.iterrows(), total=len(sample_metadata), desc="KNN MLMarker Training"):
- sample_name = row['sample']
- cohort = row['Cohort']
- expr_col = row['expression_col']
- sample_profile = sample_expr_log[expr_col]
- sample_profile = sample_profile[sample_profile > 0]
- knn_results['sample'].append(sample_name)
- knn_results['cohort'].append(cohort)
- # Get both strict and extended ground truth
- true_strict = cancer_to_mlmarker.get(cohort)
- true_extended = cancer_to_mlmarker_extended.get(cohort, [])
- for k in k_values:
- pred, top_tissues, class_proba = knn_classify_with_topk(sample_profile, ml_train_X, ml_train_y, k=k)
- knn_results[f'knn_ml_k{k}_pred'].append(pred)
- # Calculate top-k accuracy using class probabilities
- if class_proba:
- sorted_by_proba = sorted(class_proba.keys(), key=lambda x: class_proba[x], reverse=True)
- for topk in [1, 2, 3, 4, 5]:
- top_k_tissues = sorted_by_proba[:topk]
- # Strict: exact 1-to-1 mapping
- is_correct_strict = true_strict in top_k_tissues if true_strict else False
- knn_results[f'knn_ml_k{k}_top{topk}_strict'].append(is_correct_strict)
- # Extended: 1-to-many mapping
- is_correct_extended = any(t in top_k_tissues for t in true_extended)
- knn_results[f'knn_ml_k{k}_top{topk}_extended'].append(is_correct_extended)
- else:
- for topk in [1, 2, 3, 4, 5]:
- knn_results[f'knn_ml_k{k}_top{topk}_strict'].append(False)
- knn_results[f'knn_ml_k{k}_top{topk}_extended'].append(False)
- knn_ml_df = pd.DataFrame(knn_results)
- print(f"\nKNN results shape: {knn_ml_df.shape}")
- # %%
- # For HPA and PDB: Use tissue profiles as "samples" for KNN
- # This treats each tissue as a single sample
- # HPA KNN (using tissue profiles)
- # hpa_matrix is proteins × tissues, so transpose to get tissues × proteins (samples × features)
- hpa_train_X = hpa_matrix.T # Now: tissues × proteins (each tissue is a "sample")
- hpa_train_y = pd.Series(hpa_train_X.index, index=hpa_train_X.index) # Labels = tissue names
- # PDB KNN
- pdb_train_X = pdb_matrix.T # tissues × proteins
- pdb_train_y = pd.Series(pdb_train_X.index, index=pdb_train_X.index)
- print(f"HPA for KNN: {hpa_train_X.shape[0]} tissues (samples), {hpa_train_X.shape[1]} proteins (features)")
- print(f"PDB for KNN: {pdb_train_X.shape[0]} tissues (samples), {pdb_train_X.shape[1]} proteins (features)")
- # Run KNN on HPA and PDB
- knn_atlas_results = {'sample': [], 'cohort': [],
- 'knn_hpa_pred': [], 'knn_pdb_pred': []}
- # Initialize columns for BOTH strict and extended
- for topk in [1, 2, 3, 4, 5]:
- knn_atlas_results[f'knn_hpa_top{topk}_strict'] = []
- knn_atlas_results[f'knn_hpa_top{topk}_extended'] = []
- knn_atlas_results[f'knn_pdb_top{topk}_strict'] = []
- knn_atlas_results[f'knn_pdb_top{topk}_extended'] = []
- for _, row in tqdm(sample_metadata.iterrows(), total=len(sample_metadata), desc="KNN HPA/PDB"):
- sample_name = row['sample']
- cohort = row['Cohort']
- expr_col = row['expression_col']
- sample_profile = sample_expr_log[expr_col]
- sample_profile = sample_profile[sample_profile > 0]
- knn_atlas_results['sample'].append(sample_name)
- knn_atlas_results['cohort'].append(cohort)
- # HPA KNN - get both strict and extended ground truth
- true_hpa_strict = cancer_to_hpa.get(cohort)
- true_hpa_extended = cancer_to_hpa_extended.get(cohort, [])
- pred, _, class_proba = knn_classify_with_topk(sample_profile, hpa_train_X, hpa_train_y, k=1)
- knn_atlas_results['knn_hpa_pred'].append(pred)
- if class_proba:
- sorted_by_proba = sorted(class_proba.keys(), key=lambda x: class_proba[x], reverse=True)
- for topk in [1, 2, 3, 4, 5]:
- top_k_tissues = sorted_by_proba[:topk]
- is_correct_strict = true_hpa_strict in top_k_tissues if true_hpa_strict else False
- knn_atlas_results[f'knn_hpa_top{topk}_strict'].append(is_correct_strict)
- is_correct_extended = any(t in top_k_tissues for t in true_hpa_extended)
- knn_atlas_results[f'knn_hpa_top{topk}_extended'].append(is_correct_extended)
- else:
- for topk in [1, 2, 3, 4, 5]:
- knn_atlas_results[f'knn_hpa_top{topk}_strict'].append(False)
- knn_atlas_results[f'knn_hpa_top{topk}_extended'].append(False)
- # PDB KNN - get both strict and extended ground truth
- true_pdb_strict = cancer_to_pdb.get(cohort)
- true_pdb_extended = cancer_to_pdb_extended.get(cohort, [])
- pred, _, class_proba = knn_classify_with_topk(sample_profile, pdb_train_X, pdb_train_y, k=1)
- knn_atlas_results['knn_pdb_pred'].append(pred)
- if class_proba:
- sorted_by_proba = sorted(class_proba.keys(), key=lambda x: class_proba[x], reverse=True)
- for topk in [1, 2, 3, 4, 5]:
- top_k_tissues = sorted_by_proba[:topk]
- is_correct_strict = true_pdb_strict in top_k_tissues if true_pdb_strict else False
- knn_atlas_results[f'knn_pdb_top{topk}_strict'].append(is_correct_strict)
- is_correct_extended = any(t in top_k_tissues for t in true_pdb_extended)
- knn_atlas_results[f'knn_pdb_top{topk}_extended'].append(is_correct_extended)
- else:
- for topk in [1, 2, 3, 4, 5]:
- knn_atlas_results[f'knn_pdb_top{topk}_strict'].append(False)
- knn_atlas_results[f'knn_pdb_top{topk}_extended'].append(False)
- knn_atlas_df = pd.DataFrame(knn_atlas_results)
- print(f"\nKNN Atlas results shape: {knn_atlas_df.shape}")
- # %%
- # For HPA and PDB: Use tissue profiles as "samples" for KNN
- # This treats each tissue as a single sample
- # HPA KNN (using tissue profiles)
- hpa_train_X = hpa_matrix.T # tissues × proteins
- hpa_train_y = pd.Series(hpa_train_X.index, index=hpa_train_X.index)
- # PDB KNN
- pdb_train_X = pdb_matrix.T # tissues × proteins
- pdb_train_y = pd.Series(pdb_train_X.index, index=pdb_train_X.index)
- print(f"HPA for KNN: {hpa_train_X.shape[0]} tissues (samples), {hpa_train_X.shape[1]} proteins (features)")
- print(f"PDB for KNN: {pdb_train_X.shape[0]} tissues (samples), {pdb_train_X.shape[1]} proteins (features)")
- # Run KNN on HPA and PDB with multiple k values
- k_values_atlas = [1, 3, 5, 7, 11, 15, 20]
- knn_atlas_results = {'sample': [], 'cohort': []}
- # Initialize columns for all k values and both strict/extended
- for k in k_values_atlas:
- knn_atlas_results[f'knn_hpa_k{k}_pred'] = []
- knn_atlas_results[f'knn_pdb_k{k}_pred'] = []
- for topk in [1, 2, 3, 4, 5]:
- knn_atlas_results[f'knn_hpa_k{k}_top{topk}_strict'] = []
- knn_atlas_results[f'knn_hpa_k{k}_top{topk}_extended'] = []
- knn_atlas_results[f'knn_pdb_k{k}_top{topk}_strict'] = []
- knn_atlas_results[f'knn_pdb_k{k}_top{topk}_extended'] = []
- for _, row in tqdm(sample_metadata.iterrows(), total=len(sample_metadata), desc="KNN HPA/PDB"):
- sample_name = row['sample']
- cohort = row['Cohort']
- expr_col = row['expression_col']
- sample_profile = sample_expr_log[expr_col]
- sample_profile = sample_profile[sample_profile > 0]
- knn_atlas_results['sample'].append(sample_name)
- knn_atlas_results['cohort'].append(cohort)
- # Ground truth for HPA and PDB
- true_hpa_strict = cancer_to_hpa.get(cohort)
- true_hpa_extended = cancer_to_hpa_extended.get(cohort, [])
- true_pdb_strict = cancer_to_pdb.get(cohort)
- true_pdb_extended = cancer_to_pdb_extended.get(cohort, [])
- # Test different k values for HPA
- for k in k_values_atlas:
- pred_hpa, _, class_proba_hpa = knn_classify_with_topk(sample_profile, hpa_train_X, hpa_train_y, k=k)
- knn_atlas_results[f'knn_hpa_k{k}_pred'].append(pred_hpa)
- if class_proba_hpa:
- sorted_by_proba = sorted(class_proba_hpa.keys(), key=lambda x: class_proba_hpa[x], reverse=True)
- for topk in [1, 2, 3, 4, 5]:
- top_k_tissues = sorted_by_proba[:topk]
- is_correct_strict = true_hpa_strict in top_k_tissues if true_hpa_strict else False
- is_correct_extended = any(t in top_k_tissues for t in true_hpa_extended)
- knn_atlas_results[f'knn_hpa_k{k}_top{topk}_strict'].append(is_correct_strict)
- knn_atlas_results[f'knn_hpa_k{k}_top{topk}_extended'].append(is_correct_extended)
- else:
- for topk in [1, 2, 3, 4, 5]:
- knn_atlas_results[f'knn_hpa_k{k}_top{topk}_strict'].append(False)
- knn_atlas_results[f'knn_hpa_k{k}_top{topk}_extended'].append(False)
- # Test different k values for PDB
- for k in k_values_atlas:
- pred_pdb, _, class_proba_pdb = knn_classify_with_topk(sample_profile, pdb_train_X, pdb_train_y, k=k)
- knn_atlas_results[f'knn_pdb_k{k}_pred'].append(pred_pdb)
- if class_proba_pdb:
- sorted_by_proba = sorted(class_proba_pdb.keys(), key=lambda x: class_proba_pdb[x], reverse=True)
- for topk in [1, 2, 3, 4, 5]:
- top_k_tissues = sorted_by_proba[:topk]
- is_correct_strict = true_pdb_strict in top_k_tissues if true_pdb_strict else False
- is_correct_extended = any(t in top_k_tissues for t in true_pdb_extended)
- knn_atlas_results[f'knn_pdb_k{k}_top{topk}_strict'].append(is_correct_strict)
- knn_atlas_results[f'knn_pdb_k{k}_top{topk}_extended'].append(is_correct_extended)
- else:
- for topk in [1, 2, 3, 4, 5]:
- knn_atlas_results[f'knn_pdb_k{k}_top{topk}_strict'].append(False)
- knn_atlas_results[f'knn_pdb_k{k}_top{topk}_extended'].append(False)
- knn_atlas_df_multik = pd.DataFrame(knn_atlas_results)
- print(f"\nKNN Atlas results shape: {knn_atlas_df_multik.shape}")
- # %%
- # KNN Atlas: Top-1 Strict Accuracy vs k value
- fig, ax = plt.subplots(figsize=(10, 6))
- # HPA KNN - Top-1 Strict
- hpa_knn_strict_accs = [knn_atlas_df_multik[f'knn_hpa_k{k}_top5_extended'].mean() * 100 for k in k_values_atlas]
- ax.plot(k_values_atlas, hpa_knn_strict_accs, 'o-', color='#2A9D8F', linewidth=2.5,
- markersize=8, label='HPA (KNN)')
- # PDB KNN - Top-1 Strict
- pdb_knn_strict_accs = [knn_atlas_df_multik[f'knn_pdb_k{k}_top5_extended'].mean() * 100 for k in k_values_atlas]
- ax.plot(k_values_atlas, pdb_knn_strict_accs, 'D-', color='#F4A261', linewidth=2.5,
- markersize=8, label='PDB (KNN)')
- ax.set_xlabel('k (Number of Neighbors)', fontsize=12, fontweight='bold')
- ax.set_ylabel('Top-1 Strict Accuracy (%)', fontsize=12, fontweight='bold')
- ax.set_title('KNN Atlas Methods: Top-1 Strict Accuracy vs k', fontsize=14, fontweight='bold')
- ax.set_xticks(k_values_atlas)
- ax.set_ylim(0, 100)
- ax.grid(True, alpha=0.3)
- ax.legend(loc='best', fontsize=11)
- plt.tight_layout()
- # plt.savefig('knn_atlas_top5_extended_vs_k.png', dpi=300, bbox_inches='tight')
- # plt.savefig('knn_atlas_top5_extended_vs_k.svg', bbox_inches='tight')
- plt.show()
- print("KNN Atlas - Top-1 Strict Accuracy by k:")
- print("=" * 50)
- print(f"{'k':<5} {'HPA (%)':<12} {'PDB (%)':<12}")
- print("-" * 50)
- for k, hpa_acc, pdb_acc in zip(k_values_atlas, hpa_knn_strict_accs, pdb_knn_strict_accs):
- print(f"{k:<5} {hpa_acc:<11.1f} {pdb_acc:<11.1f}")
- print("-" * 50)
- # Spearman baselines computed in later cell (see cells 26-27)
- # %% [markdown]
- # ## 10. Figure 1: Top-k Accuracy Line Plot
- # %%
- # Color palette for methods - manuscript style
- # MLMarker: reddish, Training atlas: bluish, HPA: greenish, PDB: yellowish
- colors_spearman = {'MLMarker': '#E63946', 'Training Atlas': '#457B9D',
- 'HPA Atlas': '#2A9D8F', 'PDB Atlas': '#E9C46A'}
- # Extended colors for Spearman vs KNN comparison
- colors_all = {
- 'MLMarker': '#E63946',
- 'Training Atlas (Spearman)': '#457B9D',
- 'Training Atlas (KNN)': '#1D3557',
- 'HPA (Spearman)': '#2A9D8F',
- 'HPA (KNN)': '#264653',
- 'PDB (Spearman)': '#E9C46A',
- 'PDB (KNN)': '#F4A261'
- }
- # ============================================
- # Figure 1: Overall Top-k Accuracy - Spearman vs KNN Comparison
- # ============================================
- fig, axes = plt.subplots(1, 2, figsize=(14, 5))
- # Prepare data - need to recalculate with filtered cohorts
- # Filter results to only include eval_cohorts
- results_filtered = results_df[results_df['cohort'].isin(eval_cohorts)]
- knn_ml_filtered = knn_ml_df[knn_ml_df['cohort'].isin(eval_cohorts)]
- knn_atlas_filtered = knn_atlas_df[knn_atlas_df['cohort'].isin(eval_cohorts)]
- # Plot 1: Extended accuracy (Extended mapping - more meaningful)
- ax = axes[0]
- # MLMarker Model
- ml_topk_ext = [results_filtered[f'mlmarker_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_topk_ext, 'o-', color=colors_all['MLMarker'],
- linewidth=2.5, markersize=8, label='MLMarker Model')
- # Training Atlas - Spearman
- ml_train_spear = [results_filtered[f'ml_training_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train_spear, 's-', color=colors_all['Training Atlas (Spearman)'],
- linewidth=2, markersize=7, label='Training Atlas (Spearman)')
- # Training Atlas - KNN k=5
- ml_train_knn = [knn_ml_filtered[f'knn_ml_k5_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train_knn, 's-', color=colors_all['Training Atlas (KNN)'],
- linewidth=2, markersize=7, label='Training Atlas (KNN)')
- # HPA - Spearman
- hpa_spear = [results_filtered[f'hpa_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_spear, '^-', color=colors_all['HPA (Spearman)'],
- linewidth=2, markersize=7, label='HPA (Spearman)')
- # HPA - KNN
- hpa_knn = [knn_atlas_filtered[f'knn_hpa_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_knn, '^-', color=colors_all['HPA (KNN)'],
- linewidth=2, markersize=7, label='HPA (KNN)')
- # PDB - Spearman
- pdb_spear = [results_filtered[f'pdb_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_spear, 'D-', color=colors_all['PDB (Spearman)'],
- linewidth=2, markersize=7, label='PDB (Spearman)')
- # PDB - KNN
- pdb_knn = [knn_atlas_filtered[f'knn_pdb_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_knn, 'D-', color=colors_all['PDB (KNN)'],
- linewidth=2, markersize=7, label='PDB (KNN)')
- ax.set_xlabel('k (Top-k)', fontsize=12)
- ax.set_ylabel('Extended Accuracy (%)', fontsize=12)
- ax.set_title('A) Overall Top-k Accuracy', fontsize=12, fontweight='bold')
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(0, 100)
- ax.grid(True, alpha=0.3)
- # Plot 2: Top-1 accuracy bar comparison
- ax = axes[1]
- methods_bar = ['MLMarker', 'Training Atlas\n(Spearman)', 'Training Atlas\n(KNN)',
- 'HPA\n(Spearman)', 'HPA\n(KNN)', 'PDB\n(Spearman)', 'PDB\n(KNN)']
- accs_bar = [ml_topk_ext[0], ml_train_spear[0], ml_train_knn[0],
- hpa_spear[0], hpa_knn[0], pdb_spear[0], pdb_knn[0]]
- bar_colors = [colors_all['MLMarker'], colors_all['Training Atlas (Spearman)'], colors_all['Training Atlas (KNN)'],
- colors_all['HPA (Spearman)'], colors_all['HPA (KNN)'],
- colors_all['PDB (Spearman)'], colors_all['PDB (KNN)']]
- bars = ax.bar(methods_bar, accs_bar, color=bar_colors, edgecolor='black', linewidth=0.5)
- for bar in bars:
- height = bar.get_height()
- ax.annotate(f'{height:.1f}%', xy=(bar.get_x() + bar.get_width()/2, height),
- xytext=(0, 3), textcoords='offset points', ha='center', va='bottom', fontsize=9)
- ax.set_ylabel('Top-1 Extended Accuracy (%)', fontsize=12)
- ax.set_title('B) Top-1 Accuracy Comparison', fontsize=12, fontweight='bold')
- ax.set_ylim(-5, 100) # Start below 0 to see all bars clearly
- ax.tick_params(axis='x', rotation=45)
- ax.grid(True, alpha=0.3, axis='y')
- # Shared legend
- handles, labels = axes[0].get_legend_handles_labels()
- fig.legend(handles, labels, loc='center right', bbox_to_anchor=(1.18, 0.5), fontsize=9)
- plt.tight_layout()
- plt.savefig('baseline_spearman_vs_knn_overall.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_spearman_vs_knn_overall.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_spearman_vs_knn_overall.png/svg")
- # ============================================
- # Figure 2: Per-Cohort Top-k Accuracy - Spearman vs KNN
- # ============================================
- fig, axes = plt.subplots(2, len(eval_cohorts), figsize=(18, 8), sharey=True)
- for col_idx, cohort in enumerate(eval_cohorts):
- # Filter data for this cohort
- cohort_results = results_filtered[results_filtered['cohort'] == cohort]
- cohort_knn_ml = knn_ml_filtered[knn_ml_filtered['cohort'] == cohort]
- cohort_knn_atlas = knn_atlas_filtered[knn_atlas_filtered['cohort'] == cohort]
- n_samples = len(cohort_results)
- # Row 1: Spearman methods
- ax = axes[0, col_idx]
- # MLMarker Model
- ml_topk = [cohort_results[f'mlmarker_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_topk, 'o-', color=colors_all['MLMarker'], linewidth=2, markersize=6, label='MLMarker')
- # Training Atlas Spearman
- ml_train = [cohort_results[f'ml_training_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train, 's-', color=colors_all['Training Atlas (Spearman)'], linewidth=1.5, markersize=5)
- # HPA Spearman
- hpa = [cohort_results[f'hpa_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa, '^-', color=colors_all['HPA (Spearman)'], linewidth=1.5, markersize=5)
- # PDB Spearman
- pdb = [cohort_results[f'pdb_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb, 'D-', color=colors_all['PDB (Spearman)'], linewidth=1.5, markersize=5)
- ax.set_title(f'{cohort}\n(n={n_samples})', fontsize=10, fontweight='bold')
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(-5, 105) # Start below 0
- ax.grid(True, alpha=0.3)
- if col_idx == 0:
- ax.set_ylabel('Spearman (%)', fontsize=10)
- # Row 2: KNN methods
- ax = axes[1, col_idx]
- # MLMarker Model (same as above for reference)
- ax.plot([1,2,3,4,5], ml_topk, 'o-', color=colors_all['MLMarker'], linewidth=2, markersize=6)
- # Training Atlas KNN
- ml_knn = [cohort_knn_ml[f'knn_ml_k5_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_knn, 's-', color=colors_all['Training Atlas (KNN)'], linewidth=1.5, markersize=5)
- # HPA KNN
- hpa_k = [cohort_knn_atlas[f'knn_hpa_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_k, '^-', color=colors_all['HPA (KNN)'], linewidth=1.5, markersize=5)
- # PDB KNN
- pdb_k = [cohort_knn_atlas[f'knn_pdb_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_k, 'D-', color=colors_all['PDB (KNN)'], linewidth=1.5, markersize=5)
- ax.set_xlabel('k', fontsize=10)
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(-5, 105) # Start below 0
- ax.grid(True, alpha=0.3)
- if col_idx == 0:
- ax.set_ylabel('KNN (%)', fontsize=10)
- # Row labels
- axes[0, 0].annotate('Spearman\nCorrelation', xy=(-0.45, 0.5), xycoords='axes fraction',
- fontsize=11, fontweight='bold', ha='center', va='center', rotation=90)
- axes[1, 0].annotate('KNN\nClassification', xy=(-0.45, 0.5), xycoords='axes fraction',
- fontsize=11, fontweight='bold', ha='center', va='center', rotation=90)
- # Legend
- legend_labels = ['MLMarker', 'Training Atlas', 'HPA Atlas', 'PDB Atlas']
- legend_colors = [colors_all['MLMarker'], colors_all['Training Atlas (Spearman)'],
- colors_all['HPA (Spearman)'], colors_all['PDB (Spearman)']]
- handles = [plt.Line2D([0], [0], color=c, linewidth=2, marker='o', markersize=6) for c in legend_colors]
- fig.legend(handles, legend_labels, loc='lower center', bbox_to_anchor=(0.5, -0.02), ncol=4, fontsize=11)
- plt.suptitle('Top-k Accuracy by Cohort: Spearman vs KNN', fontsize=14, fontweight='bold', y=1.02)
- plt.tight_layout()
- # plt.savefig('baseline_spearman_vs_knn_per_cohort.png', dpi=300, bbox_inches='tight')
- # plt.savefig('baseline_spearman_vs_knn_per_cohort.svg', bbox_inches='tight')
- plt.show()
- # print("Saved: baseline_spearman_vs_knn_per_cohort.png/svg")
- # %%
- # ============================================
- # STRICT (1-to-1) MAPPING VERSION
- # Same figures as above but using strict mapping instead of extended
- # ============================================
- # Color palette for methods - manuscript style
- colors_all = {
- 'MLMarker': '#E63946',
- 'Training Atlas (Spearman)': '#457B9D',
- 'Training Atlas (KNN)': '#1D3557',
- 'HPA (Spearman)': '#2A9D8F',
- 'HPA (KNN)': '#264653',
- 'PDB (Spearman)': '#E9C46A',
- 'PDB (KNN)': '#F4A261'
- }
- # ============================================
- # Figure 1: Overall Top-k Accuracy - Spearman vs KNN (STRICT 1-to-1)
- # ============================================
- fig, axes = plt.subplots(1, 2, figsize=(14, 5))
- # Prepare data - need to recalculate with filtered cohorts
- results_filtered = results_df[results_df['cohort'].isin(eval_cohorts)]
- knn_ml_filtered = knn_ml_df[knn_ml_df['cohort'].isin(eval_cohorts)]
- knn_atlas_filtered = knn_atlas_df[knn_atlas_df['cohort'].isin(eval_cohorts)]
- # Plot 1: Strict accuracy (1-to-1 mapping)
- ax = axes[0]
- # MLMarker Model - STRICT
- ml_topk_strict = [results_filtered[f'mlmarker_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_topk_strict, 'o-', color=colors_all['MLMarker'],
- linewidth=2.5, markersize=8, label='MLMarker Model')
- # Training Atlas - Spearman STRICT
- ml_train_spear_strict = [results_filtered[f'ml_training_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train_spear_strict, 's-', color=colors_all['Training Atlas (Spearman)'],
- linewidth=2, markersize=7, label='Training Atlas (Spearman)')
- # Training Atlas - KNN k=5 STRICT
- ml_train_knn_strict = [knn_ml_filtered[f'knn_ml_k5_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train_knn_strict, 's-', color=colors_all['Training Atlas (KNN)'],
- linewidth=2, markersize=7, label='Training Atlas (KNN)')
- # HPA - Spearman STRICT
- hpa_spear_strict = [results_filtered[f'hpa_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_spear_strict, '^-', color=colors_all['HPA (Spearman)'],
- linewidth=2, markersize=7, label='HPA (Spearman)')
- # HPA - KNN STRICT
- hpa_knn_strict = [knn_atlas_filtered[f'knn_hpa_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_knn_strict, '^-', color=colors_all['HPA (KNN)'],
- linewidth=2, markersize=7, label='HPA (KNN)')
- # PDB - Spearman STRICT
- pdb_spear_strict = [results_filtered[f'pdb_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_spear_strict, 'D-', color=colors_all['PDB (Spearman)'],
- linewidth=2, markersize=7, label='PDB (Spearman)')
- # PDB - KNN STRICT
- pdb_knn_strict = [knn_atlas_filtered[f'knn_pdb_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_knn_strict, 'D-', color=colors_all['PDB (KNN)'],
- linewidth=2, markersize=7, label='PDB (KNN)')
- ax.set_xlabel('k (Top-k)', fontsize=12)
- ax.set_ylabel('Strict Accuracy (%)', fontsize=12)
- ax.set_title('A) Overall Top-k Accuracy (1-to-1 Mapping)', fontsize=12, fontweight='bold')
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(0, 100)
- ax.grid(True, alpha=0.3)
- # Plot 2: Top-1 accuracy bar comparison (STRICT)
- ax = axes[1]
- methods_bar = ['MLMarker', 'Training Atlas\n(Spearman)', 'Training Atlas\n(KNN)',
- 'HPA\n(Spearman)', 'HPA\n(KNN)', 'PDB\n(Spearman)', 'PDB\n(KNN)']
- accs_bar_strict = [ml_topk_strict[0], ml_train_spear_strict[0], ml_train_knn_strict[0],
- hpa_spear_strict[0], hpa_knn_strict[0], pdb_spear_strict[0], pdb_knn_strict[0]]
- bar_colors = [colors_all['MLMarker'], colors_all['Training Atlas (Spearman)'], colors_all['Training Atlas (KNN)'],
- colors_all['HPA (Spearman)'], colors_all['HPA (KNN)'],
- colors_all['PDB (Spearman)'], colors_all['PDB (KNN)']]
- bars = ax.bar(methods_bar, accs_bar_strict, color=bar_colors, edgecolor='black', linewidth=0.5)
- for bar in bars:
- height = bar.get_height()
- ax.annotate(f'{height:.1f}%', xy=(bar.get_x() + bar.get_width()/2, height),
- xytext=(0, 3), textcoords='offset points', ha='center', va='bottom', fontsize=9)
- ax.set_ylabel('Top-1 Strict Accuracy (%)', fontsize=12)
- ax.set_title('B) Top-1 Accuracy Comparison (1-to-1 Mapping)', fontsize=12, fontweight='bold')
- ax.set_ylim(-5, 100)
- ax.tick_params(axis='x', rotation=45)
- ax.grid(True, alpha=0.3, axis='y')
- # Shared legend
- handles, labels = axes[0].get_legend_handles_labels()
- fig.legend(handles, labels, loc='center right', bbox_to_anchor=(1.18, 0.5), fontsize=9)
- plt.tight_layout()
- plt.savefig('baseline_spearman_vs_knn_overall_strict.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_spearman_vs_knn_overall_strict.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_spearman_vs_knn_overall_strict.png/svg")
- # ============================================
- # Figure 2: Per-Cohort Top-k Accuracy - Spearman vs KNN (STRICT 1-to-1)
- # ============================================
- fig, axes = plt.subplots(2, len(eval_cohorts), figsize=(18, 8), sharey=True)
- for col_idx, cohort in enumerate(eval_cohorts):
- # Filter data for this cohort
- cohort_results = results_filtered[results_filtered['cohort'] == cohort]
- cohort_knn_ml = knn_ml_filtered[knn_ml_filtered['cohort'] == cohort]
- cohort_knn_atlas = knn_atlas_filtered[knn_atlas_filtered['cohort'] == cohort]
- n_samples = len(cohort_results)
- # Row 1: Spearman methods (STRICT)
- ax = axes[0, col_idx]
- # MLMarker Model
- ml_topk = [cohort_results[f'mlmarker_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_topk, 'o-', color=colors_all['MLMarker'], linewidth=2, markersize=6, label='MLMarker')
- # Training Atlas Spearman
- ml_train = [cohort_results[f'ml_training_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train, 's-', color=colors_all['Training Atlas (Spearman)'], linewidth=1.5, markersize=5)
- # HPA Spearman
- hpa = [cohort_results[f'hpa_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa, '^-', color=colors_all['HPA (Spearman)'], linewidth=1.5, markersize=5)
- # PDB Spearman
- pdb = [cohort_results[f'pdb_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb, 'D-', color=colors_all['PDB (Spearman)'], linewidth=1.5, markersize=5)
- ax.set_title(f'{cohort}\n(n={n_samples})', fontsize=10, fontweight='bold')
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(-5, 105)
- ax.grid(True, alpha=0.3)
- if col_idx == 0:
- ax.set_ylabel('Spearman (%)', fontsize=10)
- # Row 2: KNN methods (STRICT)
- ax = axes[1, col_idx]
- # MLMarker Model (same as above for reference)
- ax.plot([1,2,3,4,5], ml_topk, 'o-', color=colors_all['MLMarker'], linewidth=2, markersize=6)
- # Training Atlas KNN
- ml_knn = [cohort_knn_ml[f'knn_ml_k5_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_knn, 's-', color=colors_all['Training Atlas (KNN)'], linewidth=1.5, markersize=5)
- # HPA KNN
- hpa_k = [cohort_knn_atlas[f'knn_hpa_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_k, '^-', color=colors_all['HPA (KNN)'], linewidth=1.5, markersize=5)
- # PDB KNN
- pdb_k = [cohort_knn_atlas[f'knn_pdb_top{k}_strict'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_k, 'D-', color=colors_all['PDB (KNN)'], linewidth=1.5, markersize=5)
- ax.set_xlabel('k', fontsize=10)
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(-5, 105)
- ax.grid(True, alpha=0.3)
- if col_idx == 0:
- ax.set_ylabel('KNN (%)', fontsize=10)
- # Row labels
- axes[0, 0].annotate('Spearman\nCorrelation', xy=(-0.45, 0.5), xycoords='axes fraction',
- fontsize=11, fontweight='bold', ha='center', va='center', rotation=90)
- axes[1, 0].annotate('KNN\nClassification', xy=(-0.45, 0.5), xycoords='axes fraction',
- fontsize=11, fontweight='bold', ha='center', va='center', rotation=90)
- # Legend
- legend_labels = ['MLMarker', 'Training Atlas', 'HPA Atlas', 'PDB Atlas']
- legend_colors = [colors_all['MLMarker'], colors_all['Training Atlas (Spearman)'],
- colors_all['HPA (Spearman)'], colors_all['PDB (Spearman)']]
- handles = [plt.Line2D([0], [0], color=c, linewidth=2, marker='o', markersize=6) for c in legend_colors]
- fig.legend(handles, legend_labels, loc='lower center', bbox_to_anchor=(0.5, -0.02), ncol=4, fontsize=11)
- plt.suptitle('Top-k Accuracy by Cohort: Spearman vs KNN (1-to-1 Mapping)', fontsize=14, fontweight='bold', y=1.02)
- plt.tight_layout()
- # plt.savefig('baseline_spearman_vs_knn_per_cohort_strict.png', dpi=300, bbox_inches='tight')
- # plt.savefig('baseline_spearman_vs_knn_per_cohort_strict.svg', bbox_inches='tight')
- plt.show()
- # print("Saved: baseline_spearman_vs_knn_per_cohort_strict.png/svg")
- # %%
- # ============================================
- # FIGURE 1: Grouped Bar Chart - Top-1 Accuracy (Extended vs Strict)
- # ============================================
- import matplotlib.gridspec as gridspec
- import seaborn as sns
- from matplotlib.lines import Line2D
- from matplotlib.patches import Patch
- # Color palette - manuscript colors
- colors_atlas = {
- 'MLMarker': '#E63946',
- 'Training Atlas': '#457B9D',
- 'HPA': '#2A9D8F',
- 'PDB': '#E9C46A'
- }
- # Filter data
- results_filtered = results_df[results_df['cohort'].isin(eval_cohorts)]
- knn_ml_filtered = knn_ml_df[knn_ml_df['cohort'].isin(eval_cohorts)]
- knn_atlas_filtered = knn_atlas_df[knn_atlas_df['cohort'].isin(eval_cohorts)]
- # Calculate Top-1 accuracies for all methods (overall averages)
- # Spearman methods
- ml_ext = results_filtered['mlmarker_top1_extended'].mean() * 100
- ml_strict = results_filtered['mlmarker_top1_strict'].mean() * 100
- train_spear_ext = results_filtered['ml_training_top1_extended'].mean() * 100
- train_spear_strict = results_filtered['ml_training_top1_strict'].mean() * 100
- hpa_spear_ext = results_filtered['hpa_top1_extended'].mean() * 100
- hpa_spear_strict = results_filtered['hpa_top1_strict'].mean() * 100
- pdb_spear_ext = results_filtered['pdb_top1_extended'].mean() * 100
- pdb_spear_strict = results_filtered['pdb_top1_strict'].mean() * 100
- # KNN methods
- train_knn_ext = knn_ml_filtered['knn_ml_k5_top1_extended'].mean() * 100
- train_knn_strict = knn_ml_filtered['knn_ml_k5_top1_strict'].mean() * 100
- hpa_knn_ext = knn_atlas_filtered['knn_hpa_top1_extended'].mean() * 100
- hpa_knn_strict = knn_atlas_filtered['knn_hpa_top1_strict'].mean() * 100
- pdb_knn_ext = knn_atlas_filtered['knn_pdb_top1_extended'].mean() * 100
- pdb_knn_strict = knn_atlas_filtered['knn_pdb_top1_strict'].mean() * 100
- # Method labels and data
- method_labels_bar = ['MLMarker', 'Training\n(Spearman)', 'Training\n(KNN)', 'HPA\n(Spearman)', 'HPA\n(KNN)', 'PDB\n(Spearman)', 'PDB\n(KNN)']
- ext_values = [ml_ext, train_spear_ext, train_knn_ext, hpa_spear_ext, hpa_knn_ext, pdb_spear_ext, pdb_knn_ext]
- strict_values = [ml_strict, train_spear_strict, train_knn_strict, hpa_spear_strict, hpa_knn_strict, pdb_spear_strict, pdb_knn_strict]
- bar_colors = [colors_atlas['MLMarker'], colors_atlas['Training Atlas'], colors_atlas['Training Atlas'],
- colors_atlas['HPA'], colors_atlas['HPA'], colors_atlas['PDB'], colors_atlas['PDB']]
- fig, ax = plt.subplots(figsize=(10, 5))
- x = np.arange(len(method_labels_bar))
- width = 0.35
- bars_ext = ax.bar(x - width/2, ext_values, width, label='Extended (1-to-many)', color=bar_colors, alpha=1.0, edgecolor='black', linewidth=1)
- bars_strict = ax.bar(x + width/2, strict_values, width, label='Strict (1-to-1)', color=bar_colors, alpha=0.5, edgecolor='black', linewidth=1, hatch='//')
- ax.set_ylabel('Top-1 Accuracy (%)', fontsize=12, fontweight='bold')
- ax.set_title('Top-1 Accuracy Comparison: Extended vs Strict Mapping', fontsize=13, fontweight='bold')
- ax.set_xticks(x)
- ax.set_xticklabels(method_labels_bar, fontsize=10)
- ax.set_ylim(0, 105)
- ax.grid(True, alpha=0.3, axis='y')
- # Add value labels on bars
- for bar in bars_ext:
- height = bar.get_height()
- ax.annotate(f'{height:.1f}', xy=(bar.get_x() + bar.get_width()/2, height),
- xytext=(0, 3), textcoords="offset points", ha='center', va='bottom', fontsize=9)
- for bar in bars_strict:
- height = bar.get_height()
- ax.annotate(f'{height:.1f}', xy=(bar.get_x() + bar.get_width()/2, height),
- xytext=(0, 3), textcoords="offset points", ha='center', va='bottom', fontsize=9)
- # Custom legend for bar chart
- legend_elements = [Patch(facecolor='gray', edgecolor='black', alpha=1.0, label='Extended (1-to-many)'),
- Patch(facecolor='gray', edgecolor='black', alpha=0.5, hatch='//', label='Strict (1-to-1)')]
- ax.legend(handles=legend_elements, loc='upper right', fontsize=10)
- plt.tight_layout()
- # plt.savefig('../Figures_paper/fig11a.png', dpi=300, bbox_inches='tight')
- # plt.savefig('baseline_barplot_extended_vs_strict.svg', bbox_inches='tight')
- plt.show()
- # print("Saved: ../Figures_paper/fig11a.png/svg")
- # %%
- # ============================================
- # FIGURE 2: Per-Cohort Line Plots (Extended and Strict)
- # ============================================
- fig, axes = plt.subplots(2, len(eval_cohorts), figsize=(16, 7), sharey=True)
- row_labels = ['Strict (1-to-1)', 'Extended (1-to-many)']
- mapping_types = ['strict', 'extended']
- for row_idx, mapping_type in enumerate(mapping_types):
- for col_idx, cohort in enumerate(eval_cohorts):
- ax = axes[row_idx, col_idx]
- cohort_results = results_filtered[results_filtered['cohort'] == cohort]
- cohort_knn_ml = knn_ml_filtered[knn_ml_filtered['cohort'] == cohort]
- cohort_knn_atlas = knn_atlas_filtered[knn_atlas_filtered['cohort'] == cohort]
- n_samples = len(cohort_results)
- # MLMarker (only model - solid line)
- ml_topk = [cohort_results[f'mlmarker_top{k}_{mapping_type}'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_topk, 'o-', color=colors_atlas['MLMarker'], linewidth=2, markersize=5, label='MLMarker')
- # Training Atlas - Spearman (solid) and KNN (dashed)
- ml_train_spear = [cohort_results[f'ml_training_top{k}_{mapping_type}'].mean()*100 for k in [1,2,3,4,5]]
- ml_train_knn = [cohort_knn_ml[f'knn_ml_k5_top{k}_{mapping_type}'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train_spear, 's-', color=colors_atlas['Training Atlas'], linewidth=1.5, markersize=4, label='Training (Spearman)')
- ax.plot([1,2,3,4,5], ml_train_knn, 's--', color=colors_atlas['Training Atlas'], linewidth=1.5, markersize=4, label='Training (KNN)')
- # HPA - Spearman (solid) and KNN (dashed)
- hpa_spear = [cohort_results[f'hpa_top{k}_{mapping_type}'].mean()*100 for k in [1,2,3,4,5]]
- hpa_knn = [cohort_knn_atlas[f'knn_hpa_top{k}_{mapping_type}'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_spear, '^-', color=colors_atlas['HPA'], linewidth=1.5, markersize=4, label='HPA (Spearman)')
- ax.plot([1,2,3,4,5], hpa_knn, '^--', color=colors_atlas['HPA'], linewidth=1.5, markersize=4, label='HPA (KNN)')
- # PDB - Spearman (solid) and KNN (dashed)
- pdb_spear = [cohort_results[f'pdb_top{k}_{mapping_type}'].mean()*100 for k in [1,2,3,4,5]]
- pdb_knn = [cohort_knn_atlas[f'knn_pdb_top{k}_{mapping_type}'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_spear, 'D-', color=colors_atlas['PDB'], linewidth=1.5, markersize=4, label='PDB (Spearman)')
- ax.plot([1,2,3,4,5], pdb_knn, 'D--', color=colors_atlas['PDB'], linewidth=1.5, markersize=4, label='PDB (KNN)')
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(-5, 105)
- ax.grid(True, alpha=0.3)
- # Column titles (cohort names) only on first row
- if row_idx == 0:
- ax.set_title(f'{cohort}\n(n={n_samples})', fontsize=11, fontweight='bold')
- # Y-axis label only on first column
- if col_idx == 0:
- ax.set_ylabel(f'{row_labels[row_idx]}\nAccuracy (%)', fontsize=10, fontweight='bold')
- # X-axis label only on last row
- if row_idx == 1:
- ax.set_xlabel('k (Top-k)', fontsize=10)
- # Create legend outside the plots
- style_handles = [
- Line2D([0], [0], color='gray', linestyle='-', linewidth=2, label='Spearman'),
- Line2D([0], [0], color='gray', linestyle='--', linewidth=2, label='KNN'),
- Line2D([0], [0], marker='o', color='w', markerfacecolor=colors_atlas['MLMarker'], markersize=10, label='MLMarker'),
- Line2D([0], [0], marker='s', color='w', markerfacecolor=colors_atlas['Training Atlas'], markersize=10, label='Training Atlas'),
- Line2D([0], [0], marker='^', color='w', markerfacecolor=colors_atlas['HPA'], markersize=10, label='HPA'),
- Line2D([0], [0], marker='D', color='w', markerfacecolor=colors_atlas['PDB'], markersize=10, label='PDB'),
- ]
- fig.legend(handles=style_handles, loc='center right', bbox_to_anchor=(1.12, 0.5), fontsize=10, frameon=True, title='Method')
- plt.suptitle('Top-k Accuracy by Cohort', fontsize=14, fontweight='bold', y=1.02)
- plt.tight_layout()
- # plt.savefig('../Figures_paper/fig11c.png', dpi=300, bbox_inches='tight')
- # plt.savefig('baseline_topk_per_cohort.svg', bbox_inches='tight')
- plt.show()
- print("Saved: ../Figures_paper/fig11c.png/svg")
- # %%
- # ============================================
- # FIGURE 3: Heatmap - Top-1 Extended Accuracy by Cohort and Method
- # ============================================
- # Prepare heatmap data
- heatmap_methods = ['MLMarker', 'Training (Spearman)', 'Training (KNN)', 'HPA (Spearman)', 'HPA (KNN)', 'PDB (Spearman)', 'PDB (KNN)']
- heatmap_data = []
- for cohort in eval_cohorts:
- cohort_results = results_filtered[results_filtered['cohort'] == cohort]
- cohort_knn_ml = knn_ml_filtered[knn_ml_filtered['cohort'] == cohort]
- cohort_knn_atlas = knn_atlas_filtered[knn_atlas_filtered['cohort'] == cohort]
- row = [
- cohort_results['mlmarker_top1_extended'].mean() * 100,
- cohort_results['ml_training_top1_extended'].mean() * 100,
- cohort_knn_ml['knn_ml_k5_top1_extended'].mean() * 100,
- cohort_results['hpa_top1_extended'].mean() * 100,
- cohort_knn_atlas['knn_hpa_top1_extended'].mean() * 100,
- cohort_results['pdb_top1_extended'].mean() * 100,
- cohort_knn_atlas['knn_pdb_top1_extended'].mean() * 100,
- ]
- heatmap_data.append(row)
- heatmap_df = pd.DataFrame(heatmap_data, index=eval_cohorts, columns=heatmap_methods)
- # Colorblind-friendly colormap
- cmap = sns.color_palette('YlGn', as_cmap=True)
- fig, ax = plt.subplots(figsize=(10, 5))
- sns.heatmap(heatmap_df, annot=True, fmt='.1f', cmap=cmap,
- vmin=0, vmax=100, ax=ax, cbar_kws={'label': 'Accuracy (%)', 'shrink': 0.8},
- linewidths=0.5, linecolor='white', annot_kws={'fontsize': 11})
- ax.set_title('Top-1 Extended Accuracy by Cohort and Method', fontsize=13, fontweight='bold')
- ax.set_xlabel('Method', fontsize=12)
- ax.set_ylabel('Cohort', fontsize=12)
- ax.set_xticklabels(ax.get_xticklabels(), rotation=45, ha='right', fontsize=10)
- ax.set_yticklabels(ax.get_yticklabels(), rotation=0, fontsize=10)
- plt.tight_layout()
- # plt.savefig('../Figures_paper/fig11b.png', dpi=300, bbox_inches='tight')
- # plt.savefig('baseline_heatmap_extended.svg', bbox_inches='tight')
- plt.show()
- print("Saved: ../Figures_paper/fig11b.png/svg")
- # %%
- # %% [markdown]
- # ## 11. Figure 2: Per-Cohort Bar Plot (Top-1 Extended)
- # %%
- # Filter to Top-1 Extended
- bar_data = cohort_df[(cohort_df['k'] == 1)].copy()
- fig, ax = plt.subplots(figsize=(14, 6))
- x = np.arange(len(eval_cohorts))
- width = 0.2
- for i, method in enumerate(method_labels):
- method_data = bar_data[bar_data['Method'] == method].set_index('Cohort').loc[eval_cohorts]
- bars = ax.bar(x + i*width, method_data['Extended (%)'], width,
- label=method, color=colors[method], edgecolor='black', linewidth=0.5)
- # Add value labels
- for bar in bars:
- height = bar.get_height()
- ax.annotate(f'{height:.0f}', xy=(bar.get_x() + bar.get_width()/2, height),
- xytext=(0, 3), textcoords='offset points', ha='center', va='bottom', fontsize=8)
- ax.set_xlabel('Cancer Cohort', fontsize=12)
- ax.set_ylabel('Top-1 Extended Accuracy (%)', fontsize=12)
- ax.set_title('Classification Accuracy by Cohort (Extended Mapping)', fontsize=14, fontweight='bold')
- ax.set_xticks(x + width * 1.5)
- ax.set_xticklabels(eval_cohorts, rotation=45, ha='right')
- ax.set_ylim(0, 115)
- ax.legend(loc='upper right')
- ax.grid(True, alpha=0.3, axis='y')
- plt.tight_layout()
- plt.savefig('baseline_cohort_barplot.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_cohort_barplot.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_cohort_barplot.png/svg")
- # %% [markdown]
- # ## 12. Figure 3: Per-Cohort Top-k Line Plots
- # %%
- # Calculate full k range per cohort
- cohort_topk_data = []
- for cohort in eval_cohorts:
- cohort_mask = results_df['cohort'] == cohort
- for method, label in zip(methods, method_labels):
- for k in [1, 2, 3, 4, 5]:
- extended_acc = results_df.loc[cohort_mask, f'{method}_top{k}_extended'].mean() * 100
- cohort_topk_data.append({'Cohort': cohort, 'Method': label, 'k': k, 'Extended (%)': extended_acc})
- cohort_topk_df = pd.DataFrame(cohort_topk_data)
- # Create subplot grid
- fig, axes = plt.subplots(2, 4, figsize=(16, 8))
- axes = axes.flatten()
- for idx, cohort in enumerate(eval_cohorts):
- ax = axes[idx]
- cohort_data = cohort_topk_df[cohort_topk_df['Cohort'] == cohort]
- for method in method_labels:
- method_data = cohort_data[cohort_data['Method'] == method]
- ax.plot(method_data['k'], method_data['Extended (%)'], 'o-',
- label=method, color=colors[method], linewidth=2, markersize=6)
- n_samples = (results_df['cohort'] == cohort).sum()
- ax.set_title(f'{cohort} (n={n_samples})', fontsize=11, fontweight='bold')
- ax.set_xlabel('k')
- ax.set_ylabel('Accuracy (%)')
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(0, 105)
- ax.grid(True, alpha=0.3)
- # Hide empty subplot
- axes[-1].axis('off')
- # Add shared legend
- handles, labels = axes[0].get_legend_handles_labels()
- fig.legend(handles, labels, loc='lower right', bbox_to_anchor=(0.98, 0.12), fontsize=10)
- plt.suptitle('Top-k Accuracy by Cohort (Extended Mapping)', fontsize=14, fontweight='bold', y=1.02)
- plt.tight_layout()
- plt.savefig('baseline_cohort_topk_lines.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_cohort_topk_lines.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_cohort_topk_lines.png/svg")
- # %% [markdown]
- # ## 13. Figure 4: Accuracy Heatmap
- # %%
- # Create heatmap data (Top-1 Extended)
- heatmap_data = bar_data.pivot(index='Cohort', columns='Method', values='Extended (%)')
- heatmap_data = heatmap_data.loc[eval_cohorts, method_labels] # Reorder
- fig, ax = plt.subplots(figsize=(10, 6))
- sns.heatmap(heatmap_data, annot=True, fmt='.1f', cmap='RdYlGn',
- vmin=0, vmax=100, ax=ax, cbar_kws={'label': 'Accuracy (%)'},
- linewidths=0.5, linecolor='white')
- ax.set_title('Top-1 Extended Accuracy by Cohort and Method', fontsize=14, fontweight='bold')
- ax.set_xlabel('Method', fontsize=12)
- ax.set_ylabel('Cohort', fontsize=12)
- plt.xticks(rotation=45, ha='right')
- plt.tight_layout()
- plt.savefig('baseline_accuracy_heatmap.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_accuracy_heatmap.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_accuracy_heatmap.png/svg")
- # %% [markdown]
- # ## 14. Confusion Matrices
- # %%
- from sklearn.metrics import confusion_matrix
- def plot_confusion_matrix(y_true, y_pred, title, ax, normalize=True):
- """Plot confusion matrix with cohort labels."""
- labels = sorted(set(y_true) | set(y_pred))
- cm = confusion_matrix(y_true, y_pred, labels=labels)
- if normalize:
- cm = cm.astype('float') / cm.sum(axis=1)[:, np.newaxis]
- cm = np.nan_to_num(cm) # Handle division by zero
- sns.heatmap(cm, annot=False, fmt='.2f' if normalize else 'd', cmap='Blues',
- xticklabels=labels, yticklabels=labels, ax=ax,
- vmin=0, vmax=1 if normalize else None)
- ax.set_title(title, fontsize=11, fontweight='bold')
- ax.set_xlabel('Predicted')
- ax.set_ylabel('True Cohort')
- # Create confusion matrices for each method
- fig, axes = plt.subplots(2, 2, figsize=(16, 14))
- # MLMarker: Use predicted tissue grouped by cohort
- ax = axes[0, 0]
- y_true = results_df['cohort'].values
- y_pred = results_df['mlmarker_pred'].fillna('Unknown').values
- plot_confusion_matrix(y_true, y_pred, 'MLMarker Model', ax)
- ax = axes[0, 1]
- y_pred = results_df['ml_training_pred'].fillna('Unknown').values
- plot_confusion_matrix(y_true, y_pred, 'MLMarker Training Atlas', ax)
- ax = axes[1, 0]
- y_pred = results_df['hpa_pred'].fillna('Unknown').values
- plot_confusion_matrix(y_true, y_pred, 'HPA Atlas', ax)
- ax = axes[1, 1]
- y_pred = results_df['pdb_pred'].fillna('Unknown').values
- plot_confusion_matrix(y_true, y_pred, 'PDB Atlas', ax)
- plt.suptitle('Confusion Matrices: True Cohort vs Predicted Tissue', fontsize=14, fontweight='bold')
- plt.tight_layout()
- plt.savefig('baseline_confusion_matrices.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_confusion_matrices.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_confusion_matrices.png/svg")
- # %% [markdown]
- # ## 15. Summary Statistics and Save Results
- # %%
- # Final summary
- print("=" * 80)
- print("FINAL SUMMARY: MLMarker vs Atlas-Correlation Baselines")
- print("=" * 80)
- mlmarker_t1 = results_df['mlmarker_top1_extended'].mean() * 100
- mltraining_t1 = results_df['ml_training_top1_extended'].mean() * 100
- hpa_t1 = results_df['hpa_top1_extended'].mean() * 100
- pdb_t1 = results_df['pdb_top1_extended'].mean() * 100
- print(f"\nTop-1 Extended Accuracy:")
- print(f" MLMarker Model: {mlmarker_t1:.1f}%")
- print(f" MLMarker Training: {mltraining_t1:.1f}% (Δ = {mlmarker_t1 - mltraining_t1:+.1f}%)")
- print(f" HPA Atlas: {hpa_t1:.1f}% (Δ = {mlmarker_t1 - hpa_t1:+.1f}%)")
- print(f" PDB Atlas: {pdb_t1:.1f}% (Δ = {mlmarker_t1 - pdb_t1:+.1f}%)")
- print(f"\nEvaluated on {len(results_df)} samples across {len(eval_cohorts)} cohorts")
- print(f"Cohorts: {', '.join(eval_cohorts)}")
- # Save results
- summary_df.to_csv('baseline_comparison_summary.csv', index=False)
- cohort_df.to_csv('baseline_comparison_cohort.csv', index=False)
- results_df.to_csv('baseline_comparison_detailed.csv', index=False)
- print("\nResults saved:")
- print(" - baseline_comparison_summary.csv")
- print(" - baseline_comparison_cohort.csv")
- print(" - baseline_comparison_detailed.csv")
- # %% [markdown]
- # ## 16. Additional Baseline Methods
- #
- # ### 16.1 Nearest Centroid Classification
- # Classify samples by Euclidean distance to tissue centroids.
- # %%
- from scipy.spatial.distance import euclidean, cdist
- from sklearn.neighbors import KNeighborsClassifier
- from sklearn.preprocessing import StandardScaler
- def nearest_centroid_classify(sample_profile, atlas_matrix):
- """Classify by Euclidean distance to tissue centroids."""
- common_proteins = sample_profile.index.intersection(atlas_matrix.index)
- if len(common_proteins) < 10:
- return None, {}, {}
- sample_aligned = sample_profile.loc[common_proteins].values
- atlas_aligned = atlas_matrix.loc[common_proteins]
- distances = {}
- for tissue in atlas_aligned.columns:
- tissue_profile = atlas_aligned[tissue].values
- dist = euclidean(sample_aligned, tissue_profile)
- distances[tissue] = -dist # Negative so higher = closer
- if distances:
- predicted = max(distances, key=distances.get)
- return predicted, distances[predicted], distances
- return None, 0, {}
- # Run Nearest Centroid classification on all atlases
- nc_results = {'sample': [], 'cohort': [],
- 'nc_ml_pred': [], 'nc_hpa_pred': [], 'nc_pdb_pred': []}
- for _, row in tqdm(sample_metadata.iterrows(), total=len(sample_metadata), desc="Nearest Centroid"):
- sample_name = row['sample']
- cohort = row['Cohort']
- expr_col = row['expression_col']
- sample_profile = sample_expr_log[expr_col]
- sample_profile = sample_profile[sample_profile > 0]
- nc_results['sample'].append(sample_name)
- nc_results['cohort'].append(cohort)
- # MLMarker Training Atlas
- pred, _, scores = nearest_centroid_classify(sample_profile, ml_atlas)
- nc_results['nc_ml_pred'].append(pred)
- # HPA Atlas
- pred, _, scores = nearest_centroid_classify(sample_profile, hpa_matrix)
- nc_results['nc_hpa_pred'].append(pred)
- # PDB Atlas
- pred, _, scores = nearest_centroid_classify(sample_profile, pdb_matrix)
- nc_results['nc_pdb_pred'].append(pred)
- nc_df = pd.DataFrame(nc_results)
- # Calculate accuracy
- print("Nearest Centroid Classification Results (Top-1 Extended):")
- print("=" * 60)
- # MLMarker Training
- correct = sum(nc_df.apply(lambda r: r['nc_ml_pred'] in cancer_to_mlmarker_extended.get(r['cohort'], []), axis=1))
- print(f"MLMarker Training Atlas (NC): {correct/len(nc_df)*100:.1f}%")
- # HPA
- correct = sum(nc_df.apply(lambda r: r['nc_hpa_pred'] in cancer_to_hpa_extended.get(r['cohort'], []), axis=1))
- print(f"HPA Atlas (NC): {correct/len(nc_df)*100:.1f}%")
- # PDB
- correct = sum(nc_df.apply(lambda r: r['nc_pdb_pred'] in cancer_to_pdb_extended.get(r['cohort'], []), axis=1))
- print(f"PDB Atlas (NC): {correct/len(nc_df)*100:.1f}%")
- # %%
- # Figure: KNN vs Correlation Comparison
- fig, axes = plt.subplots(1, 2, figsize=(14, 5))
- # Plot 1: MLMarker Training - KNN k sensitivity
- ax = axes[0]
- knn_accs = [knn_ml_df[f'knn_ml_k{k}_top1_extended'].mean()*100 for k in k_values]
- spearman_acc = results_df['ml_training_top1_extended'].mean()*100
- mlmarker_acc = results_df['mlmarker_top1_extended'].mean()*100
- ax.plot(k_values, knn_accs, 'o-', color='#457B9D', linewidth=2, markersize=8, label='KNN')
- ax.axhline(y=spearman_acc, color='#2A9D8F', linestyle='--', linewidth=2, label='Spearman Correlation')
- ax.axhline(y=mlmarker_acc, color='#E63946', linestyle='-', linewidth=2, label='MLMarker Model')
- ax.set_xlabel('k (Number of Neighbors)', fontsize=12)
- ax.set_ylabel('Top-1 Extended Accuracy (%)', fontsize=12)
- ax.set_title('MLMarker Training Data: KNN vs Correlation', fontsize=14, fontweight='bold')
- ax.set_xticks(k_values)
- ax.set_ylim(0, 100)
- ax.legend(loc='lower right')
- ax.grid(True, alpha=0.3)
- # Plot 2: Comparison across all methods
- ax = axes[1]
- methods_compare = ['MLMarker\nModel', 'ML Train\n(Spearman)', 'ML Train\n(KNN k=5)',
- 'HPA\n(Spearman)', 'HPA\n(KNN)', 'PDB\n(Spearman)', 'PDB\n(KNN)']
- accs_compare = [
- mlmarker_acc,
- spearman_acc,
- knn_ml_df['knn_ml_k5_top1_extended'].mean()*100,
- results_df['hpa_top1_extended'].mean()*100,
- knn_atlas_df['knn_hpa_top1_extended'].mean()*100,
- results_df['pdb_top1_extended'].mean()*100,
- knn_atlas_df['knn_pdb_top1_extended'].mean()*100
- ]
- colors_compare = ['#E63946', '#457B9D', '#1D3557', '#2A9D8F', '#264653', '#E9C46A', '#F4A261']
- bars = ax.bar(methods_compare, accs_compare, color=colors_compare, edgecolor='black', linewidth=0.5)
- for bar in bars:
- height = bar.get_height()
- ax.annotate(f'{height:.1f}%', xy=(bar.get_x() + bar.get_width()/2, height),
- xytext=(0, 3), textcoords='offset points', ha='center', va='bottom', fontsize=9)
- ax.set_ylabel('Top-1 Extended Accuracy (%)', fontsize=12)
- ax.set_title('KNN vs Spearman Correlation', fontsize=14, fontweight='bold')
- ax.set_ylim(0, 100)
- ax.tick_params(axis='x', rotation=45)
- ax.grid(True, alpha=0.3, axis='y')
- plt.tight_layout()
- plt.savefig('baseline_knn_comparison.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_knn_comparison.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_knn_comparison.png/svg")
- # %%
- # Per-cohort KNN accuracy breakdown
- print("\nPer-Cohort KNN Accuracy (Top-1 Extended):")
- print("=" * 100)
- print(f"{'Cohort':<12} {'MLMarker':>10} {'ML Spear':>10} {'ML KNN-5':>10} {'HPA Spear':>10} {'HPA KNN':>10} {'PDB Spear':>10} {'PDB KNN':>10}")
- print("-" * 100)
- for cohort in eval_cohorts:
- mask = results_df['cohort'] == cohort
- knn_mask = knn_ml_df['cohort'] == cohort
- atlas_mask = knn_atlas_df['cohort'] == cohort
- ml_model = results_df.loc[mask, 'mlmarker_top1_extended'].mean() * 100
- ml_spear = results_df.loc[mask, 'ml_training_top1_extended'].mean() * 100
- ml_knn = knn_ml_df.loc[knn_mask, 'knn_ml_k5_top1_extended'].mean() * 100
- hpa_spear = results_df.loc[mask, 'hpa_top1_extended'].mean() * 100
- hpa_knn = knn_atlas_df.loc[atlas_mask, 'knn_hpa_top1_extended'].mean() * 100
- pdb_spear = results_df.loc[mask, 'pdb_top1_extended'].mean() * 100
- pdb_knn = knn_atlas_df.loc[atlas_mask, 'knn_pdb_top1_extended'].mean() * 100
- print(f"{cohort:<12} {ml_model:>9.1f}% {ml_spear:>9.1f}% {ml_knn:>9.1f}% {hpa_spear:>9.1f}% {hpa_knn:>9.1f}% {pdb_spear:>9.1f}% {pdb_knn:>9.1f}%")
- # %%
- # Top-k accuracy curves for KNN methods
- fig, ax = plt.subplots(figsize=(10, 6))
- # MLMarker Model
- ml_topk = [results_df[f'mlmarker_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_topk, 'o-', color='#E63946', linewidth=2, markersize=8, label='MLMarker Model')
- # ML Training Spearman
- ml_train_topk = [results_df[f'ml_training_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_train_topk, 's--', color='#457B9D', linewidth=2, markersize=8, label='ML Training (Spearman)')
- # ML Training KNN k=5
- ml_knn_topk = [knn_ml_df[f'knn_ml_k5_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], ml_knn_topk, '^--', color='#1D3557', linewidth=2, markersize=8, label='ML Training (KNN k=5)')
- # HPA Spearman
- hpa_topk = [results_df[f'hpa_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_topk, 's--', color='#2A9D8F', linewidth=2, markersize=8, label='HPA (Spearman)')
- # HPA KNN
- hpa_knn_topk = [knn_atlas_df[f'knn_hpa_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], hpa_knn_topk, '^--', color='#264653', linewidth=2, markersize=8, label='HPA (KNN)')
- # PDB Spearman
- pdb_topk = [results_df[f'pdb_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_topk, 's--', color='#E9C46A', linewidth=2, markersize=8, label='PDB (Spearman)')
- # PDB KNN
- pdb_knn_topk = [knn_atlas_df[f'knn_pdb_top{k}_extended'].mean()*100 for k in [1,2,3,4,5]]
- ax.plot([1,2,3,4,5], pdb_knn_topk, '^--', color='#F4A261', linewidth=2, markersize=8, label='PDB (KNN)')
- ax.set_xlabel('k (Top-k)', fontsize=12)
- ax.set_ylabel('Extended Accuracy (%)', fontsize=12)
- ax.set_title('Top-k Accuracy: All Methods Including KNN', fontsize=14, fontweight='bold')
- ax.set_xticks([1, 2, 3, 4, 5])
- ax.set_ylim(0, 100)
- ax.legend(loc='lower right', fontsize=9)
- ax.grid(True, alpha=0.3)
- plt.tight_layout()
- # plt.savefig('baseline_topk_with_knn.png', dpi=300, bbox_inches='tight')
- # plt.savefig('baseline_topk_with_knn.svg', bbox_inches='tight')
- plt.show()
- # print("Saved: baseline_topk_with_knn.png/svg")
- # %%
- # Save KNN results
- knn_ml_df.to_csv('baseline_knn_mltraining.csv', index=False)
- knn_atlas_df.to_csv('baseline_knn_atlas.csv', index=False)
- print("KNN results saved:")
- print(" - baseline_knn_mltraining.csv")
- print(" - baseline_knn_atlas.csv")
- # %% [markdown]
- # ### 16.2 UMAP Visualization with Atlas Centroids
- #
- # Project cancer samples and atlas tissue centroids into a shared 2D space to visualize clustering.
- # %%
- import umap
- # Use MLMarker Training Atlas for UMAP (most comparable to MLMarker model)
- # Get common proteins between samples and atlas
- common_proteins = sample_expr_log.index.intersection(ml_atlas.index)
- print(f"Common proteins for UMAP: {len(common_proteins)}")
- # Prepare sample data
- sample_data = sample_expr_log.loc[common_proteins, valid_cols].T
- sample_data.index = sample_metadata['sample'].values
- # Prepare atlas centroids
- atlas_centroids = ml_atlas.loc[common_proteins].T
- # Combine for joint embedding
- combined_data = pd.concat([sample_data, atlas_centroids])
- combined_labels = list(sample_metadata['Cohort'].values) + list(atlas_centroids.index)
- is_centroid = [False] * len(sample_data) + [True] * len(atlas_centroids)
- # Fill NaN with 0 and scale
- combined_filled = combined_data.fillna(0)
- scaler = StandardScaler()
- combined_scaled = scaler.fit_transform(combined_filled)
- # Run UMAP
- print("Running UMAP...")
- reducer = umap.UMAP(n_neighbors=15, min_dist=0.1, random_state=42, n_components=2)
- embedding = reducer.fit_transform(combined_scaled)
- # Create embedding DataFrame
- umap_df = pd.DataFrame({
- 'UMAP1': embedding[:, 0],
- 'UMAP2': embedding[:, 1],
- 'Label': combined_labels,
- 'Is_Centroid': is_centroid
- })
- print(f"UMAP embedding complete: {umap_df.shape}")
- # %%
- # Figure 5: UMAP with samples and expected tissue centroids highlighted
- fig, ax = plt.subplots(figsize=(14, 10))
- # Define colors for cohorts
- cohort_colors = {'CRC': '#E63946', 'DLBCL': '#457B9D', 'DLBCL+': '#1D3557',
- 'Glioma': '#2A9D8F', 'OSCC': '#E9C46A',
- 'healthy LN': '#F4A261', 'healthy OE': '#264653'}
- # Plot samples (smaller, transparent)
- samples_df = umap_df[~umap_df['Is_Centroid']]
- for cohort in eval_cohorts:
- mask = samples_df['Label'] == cohort
- ax.scatter(samples_df.loc[mask, 'UMAP1'], samples_df.loc[mask, 'UMAP2'],
- c=cohort_colors[cohort], alpha=0.6, s=30, label=f'{cohort} samples')
- # Plot expected tissue centroids with stars
- expected_tissues = set()
- for cohort in eval_cohorts:
- expected_tissues.update(cancer_to_mlmarker_extended.get(cohort, []))
- centroids_df = umap_df[umap_df['Is_Centroid']]
- for tissue in expected_tissues:
- if tissue in centroids_df['Label'].values:
- row = centroids_df[centroids_df['Label'] == tissue]
- ax.scatter(row['UMAP1'], row['UMAP2'], marker='*', s=500,
- c='black', edgecolors='white', linewidth=2, zorder=10)
- ax.annotate(tissue, (row['UMAP1'].values[0], row['UMAP2'].values[0]),
- fontsize=9, fontweight='bold',
- xytext=(5, 5), textcoords='offset points')
- # Plot other centroids (grey, smaller)
- other_tissues = set(centroids_df['Label']) - expected_tissues
- for tissue in other_tissues:
- row = centroids_df[centroids_df['Label'] == tissue]
- ax.scatter(row['UMAP1'], row['UMAP2'], marker='*', s=150,
- c='lightgray', edgecolors='gray', linewidth=1, zorder=5)
- ax.set_xlabel('UMAP 1', fontsize=12)
- ax.set_ylabel('UMAP 2', fontsize=12)
- ax.set_title('UMAP: Cancer Samples with MLMarker Training Atlas Tissue Centroids',
- fontsize=14, fontweight='bold')
- ax.legend(loc='upper left', bbox_to_anchor=(1.02, 1), fontsize=9)
- plt.tight_layout()
- plt.savefig('baseline_umap_centroids.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_umap_centroids.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_umap_centroids.png/svg")
- # %% [markdown]
- # ### 16.3 Prediction Confidence Analysis
- #
- # Compare the "confidence margin" of correct predictions across methods. A higher margin indicates the model is more certain about its prediction.
- # %%
- # Calculate confidence margins (top1 - top2 score) for each method
- confidence_data = []
- for _, row in tqdm(sample_metadata.iterrows(), total=len(sample_metadata), desc="Confidence analysis"):
- sample_name = row['sample']
- cohort = row['Cohort']
- expr_col = row['expression_col']
- sample_profile = sample_expr_log[expr_col]
- sample_profile = sample_profile[sample_profile > 0]
- # MLMarker Model
- if sample_name in sample_mapping_mlmarker:
- ml_col = sample_mapping_mlmarker[sample_name]
- scores = mlmarker_predictions[ml_col].sort_values(ascending=False)
- margin = scores.iloc[0] - scores.iloc[1] if len(scores) > 1 else 0
- is_correct = scores.index[0] in cancer_to_mlmarker_extended.get(cohort, [])
- confidence_data.append({'Method': 'MLMarker', 'Margin': margin,
- 'Correct': is_correct, 'Cohort': cohort})
- # MLMarker Training Atlas (Spearman)
- _, _, scores = atlas_correlation_classify(sample_profile, ml_atlas)
- if scores:
- sorted_scores = sorted(scores.values(), reverse=True)
- margin = sorted_scores[0] - sorted_scores[1] if len(sorted_scores) > 1 else 0
- pred = max(scores, key=scores.get)
- is_correct = pred in cancer_to_mlmarker_extended.get(cohort, [])
- confidence_data.append({'Method': 'MLMarker Training', 'Margin': margin,
- 'Correct': is_correct, 'Cohort': cohort})
- # HPA Atlas
- _, _, scores = atlas_correlation_classify(sample_profile, hpa_matrix)
- if scores:
- sorted_scores = sorted(scores.values(), reverse=True)
- margin = sorted_scores[0] - sorted_scores[1] if len(sorted_scores) > 1 else 0
- pred = max(scores, key=scores.get)
- is_correct = pred in cancer_to_hpa_extended.get(cohort, [])
- confidence_data.append({'Method': 'HPA Atlas', 'Margin': margin,
- 'Correct': is_correct, 'Cohort': cohort})
- # PDB Atlas
- _, _, scores = atlas_correlation_classify(sample_profile, pdb_matrix)
- if scores:
- sorted_scores = sorted(scores.values(), reverse=True)
- margin = sorted_scores[0] - sorted_scores[1] if len(sorted_scores) > 1 else 0
- pred = max(scores, key=scores.get)
- is_correct = pred in cancer_to_pdb_extended.get(cohort, [])
- confidence_data.append({'Method': 'PDB Atlas', 'Margin': margin,
- 'Correct': is_correct, 'Cohort': cohort})
- confidence_df = pd.DataFrame(confidence_data)
- print(f"Confidence data collected: {len(confidence_df)} predictions")
- # %%
- # Figure 6: Confidence Margin Distribution
- fig, axes = plt.subplots(1, 2, figsize=(14, 5))
- # Plot 1: Box plot of margins by method
- ax = axes[0]
- order = ['MLMarker', 'MLMarker Training', 'HPA Atlas', 'PDB Atlas']
- palette = {'MLMarker': '#E63946', 'MLMarker Training': '#457B9D',
- 'HPA Atlas': '#2A9D8F', 'PDB Atlas': '#E9C46A'}
- sns.boxplot(data=confidence_df, x='Method', y='Margin', order=order,
- palette=palette, ax=ax)
- ax.set_xlabel('Method', fontsize=12)
- ax.set_ylabel('Confidence Margin (Top1 - Top2 Score)', fontsize=12)
- ax.set_title('Prediction Confidence by Method', fontsize=14, fontweight='bold')
- ax.tick_params(axis='x', rotation=45)
- # Plot 2: Margin distribution for correct vs incorrect predictions
- ax = axes[1]
- correct_df = confidence_df[confidence_df['Correct']]
- incorrect_df = confidence_df[~confidence_df['Correct']]
- for i, method in enumerate(order):
- correct_margins = correct_df[correct_df['Method'] == method]['Margin']
- incorrect_margins = incorrect_df[incorrect_df['Method'] == method]['Margin']
- positions = [i - 0.2, i + 0.2]
- bp = ax.boxplot([correct_margins, incorrect_margins], positions=positions,
- widths=0.35, patch_artist=True)
- bp['boxes'][0].set_facecolor(palette[method])
- bp['boxes'][0].set_alpha(0.8)
- bp['boxes'][1].set_facecolor(palette[method])
- bp['boxes'][1].set_alpha(0.3)
- ax.set_xticks(range(len(order)))
- ax.set_xticklabels(order, rotation=45, ha='right')
- ax.set_xlabel('Method', fontsize=12)
- ax.set_ylabel('Confidence Margin', fontsize=12)
- ax.set_title('Confidence: Correct (dark) vs Incorrect (light)', fontsize=14, fontweight='bold')
- # Add legend
- from matplotlib.patches import Patch
- legend_elements = [Patch(facecolor='gray', alpha=0.8, label='Correct'),
- Patch(facecolor='gray', alpha=0.3, label='Incorrect')]
- ax.legend(handles=legend_elements, loc='upper right')
- plt.tight_layout()
- plt.savefig('baseline_confidence_margins.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_confidence_margins.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_confidence_margins.png/svg")
- # Print mean margins
- print("\nMean Confidence Margins:")
- print("=" * 50)
- for method in order:
- method_df = confidence_df[confidence_df['Method'] == method]
- correct_margin = method_df[method_df['Correct']]['Margin'].mean()
- incorrect_margin = method_df[~method_df['Correct']]['Margin'].mean()
- print(f"{method:<20}: Correct={correct_margin:.3f}, Incorrect={incorrect_margin:.3f}")
- # %% [markdown]
- # ### 16.4 Statistical Significance: McNemar's Test
- #
- # Compare MLMarker vs each baseline using McNemar's test for paired nominal data.
- # %%
- from statsmodels.stats.contingency_tables import mcnemar
- def run_mcnemar(correct_a, correct_b, name_a, name_b):
- """Run McNemar's test comparing two classifiers."""
- # Create contingency table
- both_correct = sum(a and b for a, b in zip(correct_a, correct_b))
- a_only = sum(a and not b for a, b in zip(correct_a, correct_b))
- b_only = sum(not a and b for a, b in zip(correct_a, correct_b))
- both_wrong = sum(not a and not b for a, b in zip(correct_a, correct_b))
- table = [[both_correct, a_only], [b_only, both_wrong]]
- result = mcnemar(table, exact=True)
- return {
- 'Comparison': f'{name_a} vs {name_b}',
- 'Both Correct': both_correct,
- f'{name_a} Only': a_only,
- f'{name_b} Only': b_only,
- 'Both Wrong': both_wrong,
- 'p-value': result.pvalue,
- 'Significant': 'Yes' if result.pvalue < 0.05 else 'No'
- }
- # Get correct/incorrect arrays
- mlmarker_correct = results_df['mlmarker_top1_extended'].values
- ml_training_correct = results_df['ml_training_top1_extended'].values
- hpa_correct = results_df['hpa_top1_extended'].values
- pdb_correct = results_df['pdb_top1_extended'].values
- # Run McNemar's tests
- mcnemar_results = []
- mcnemar_results.append(run_mcnemar(mlmarker_correct, ml_training_correct, 'MLMarker', 'ML Training'))
- mcnemar_results.append(run_mcnemar(mlmarker_correct, hpa_correct, 'MLMarker', 'HPA'))
- mcnemar_results.append(run_mcnemar(mlmarker_correct, pdb_correct, 'MLMarker', 'PDB'))
- mcnemar_results.append(run_mcnemar(ml_training_correct, hpa_correct, 'ML Training', 'HPA'))
- mcnemar_results.append(run_mcnemar(ml_training_correct, pdb_correct, 'ML Training', 'PDB'))
- mcnemar_results.append(run_mcnemar(hpa_correct, pdb_correct, 'HPA', 'PDB'))
- mcnemar_df = pd.DataFrame(mcnemar_results)
- print("McNemar's Test Results (Top-1 Extended Accuracy):")
- print("=" * 80)
- print(mcnemar_df.to_string(index=False))
- # %% [markdown]
- # ### 16.5 Random Baseline Comparison
- #
- # Add a random classifier to show all methods perform significantly above chance.
- # %%
- # Calculate random baseline (chance level)
- # For MLMarker: 34 tissue classes
- n_mlmarker_classes = len(mlmarker_predictions.index)
- random_mlmarker_strict = 1 / n_mlmarker_classes * 100
- # Extended: average number of acceptable tissues per cohort
- avg_extended = np.mean([len(v) for v in cancer_to_mlmarker_extended.values()])
- random_mlmarker_extended = avg_extended / n_mlmarker_classes * 100
- # For HPA: 61 tissues
- n_hpa_classes = len(hpa_matrix.columns)
- avg_hpa_extended = np.mean([len(v) for v in cancer_to_hpa_extended.values()])
- random_hpa_extended = avg_hpa_extended / n_hpa_classes * 100
- # For PDB: 57 tissues
- n_pdb_classes = len(pdb_matrix.columns)
- avg_pdb_extended = np.mean([len(v) for v in cancer_to_pdb_extended.values()])
- random_pdb_extended = avg_pdb_extended / n_pdb_classes * 100
- print("Random Baseline (Chance Level):")
- print("=" * 60)
- print(f"MLMarker ({n_mlmarker_classes} classes):")
- print(f" Strict: {random_mlmarker_strict:.2f}%")
- print(f" Extended (avg {avg_extended:.1f} acceptable): {random_mlmarker_extended:.2f}%")
- print(f"\nHPA ({n_hpa_classes} tissues):")
- print(f" Extended (avg {avg_hpa_extended:.1f} acceptable): {random_hpa_extended:.2f}%")
- print(f"\nPDB ({n_pdb_classes} tissues):")
- print(f" Extended (avg {avg_pdb_extended:.1f} acceptable): {random_pdb_extended:.2f}%")
- # %% [markdown]
- # ## 17. Final Summary Figure: Combined Comparison
- # %%
- # Figure 7: Comprehensive comparison bar chart with random baseline
- fig, ax = plt.subplots(figsize=(12, 7))
- # Collect all method accuracies (Top-1 Extended)
- method_accs = {
- 'MLMarker\nModel': results_df['mlmarker_top1_extended'].mean() * 100,
- 'MLMarker\nTraining Atlas': results_df['ml_training_top1_extended'].mean() * 100,
- 'HPA Atlas\n(Spearman)': results_df['hpa_top1_extended'].mean() * 100,
- 'PDB Atlas\n(Spearman)': results_df['pdb_top1_extended'].mean() * 100,
- }
- # Add random baselines
- method_accs['Random\n(MLMarker)'] = random_mlmarker_extended
- method_accs['Random\n(HPA)'] = random_hpa_extended
- method_accs['Random\n(PDB)'] = random_pdb_extended
- # Colors
- bar_colors = ['#E63946', '#457B9D', '#2A9D8F', '#E9C46A', '#CCCCCC', '#CCCCCC', '#CCCCCC']
- x = np.arange(len(method_accs))
- bars = ax.bar(x, list(method_accs.values()), color=bar_colors, edgecolor='black', linewidth=0.5)
- # Add value labels
- for bar in bars:
- height = bar.get_height()
- ax.annotate(f'{height:.1f}%', xy=(bar.get_x() + bar.get_width()/2, height),
- xytext=(0, 3), textcoords='offset points', ha='center', va='bottom',
- fontsize=11, fontweight='bold')
- ax.set_xticks(x)
- ax.set_xticklabels(list(method_accs.keys()), fontsize=10)
- ax.set_ylabel('Top-1 Extended Accuracy (%)', fontsize=12)
- ax.set_title('Tissue Classification Accuracy: MLMarker vs Atlas Baselines',
- fontsize=14, fontweight='bold')
- ax.set_ylim(0, 100)
- ax.axhline(y=50, color='gray', linestyle='--', alpha=0.5, label='50% threshold')
- ax.grid(True, alpha=0.3, axis='y')
- # Add annotation for improvement
- best_baseline = max(method_accs['MLMarker\nTraining Atlas'],
- method_accs['HPA Atlas\n(Spearman)'],
- method_accs['PDB Atlas\n(Spearman)'])
- improvement = method_accs['MLMarker\nModel'] - best_baseline
- ax.annotate(f'+{improvement:.1f}% vs best baseline',
- xy=(0, method_accs['MLMarker\nModel']),
- xytext=(1.5, method_accs['MLMarker\nModel'] + 5),
- fontsize=11, fontweight='bold', color='#E63946',
- arrowprops=dict(arrowstyle='->', color='#E63946'))
- plt.tight_layout()
- plt.savefig('baseline_final_comparison.png', dpi=300, bbox_inches='tight')
- plt.savefig('baseline_final_comparison.svg', bbox_inches='tight')
- plt.show()
- print("Saved: baseline_final_comparison.png/svg")
- # %%
- # Save all results
- mcnemar_df.to_csv('baseline_mcnemar_tests.csv', index=False)
- confidence_df.to_csv('baseline_confidence_margins.csv', index=False)
- print("\n" + "=" * 80)
- print("COMPLETE ANALYSIS SUMMARY")
- print("=" * 80)
- print(f"\nDataset: {len(results_df)} cancer samples across {len(eval_cohorts)} cohorts")
- print(f"Cohorts: {', '.join(eval_cohorts)}")
- print(f"\nMethods compared:")
- print(f" 1. MLMarker Model (Random Forest, 34 tissues)")
- print(f" 2. MLMarker Training Atlas (Spearman correlation)")
- print(f" 3. HPA Atlas (Spearman correlation, {len(hpa_matrix.columns)} tissues)")
- print(f" 4. PDB Atlas (Spearman correlation, {len(pdb_matrix.columns)} tissues)")
- print(f" 5. Nearest Centroid (Euclidean distance)")
- print(f" 6. Random baseline (chance level)")
- print(f"\n{'Method':<25} {'Top-1 Extended':>15} {'vs Random':>12}")
- print("-" * 55)
- print(f"{'MLMarker Model':<25} {results_df['mlmarker_top1_extended'].mean()*100:>14.1f}% {'+' + str(round(results_df['mlmarker_top1_extended'].mean()*100 - random_mlmarker_extended, 1)):>11}%")
- print(f"{'MLMarker Training':<25} {results_df['ml_training_top1_extended'].mean()*100:>14.1f}% {'+' + str(round(results_df['ml_training_top1_extended'].mean()*100 - random_mlmarker_extended, 1)):>11}%")
- print(f"{'HPA Atlas':<25} {results_df['hpa_top1_extended'].mean()*100:>14.1f}% {'+' + str(round(results_df['hpa_top1_extended'].mean()*100 - random_hpa_extended, 1)):>11}%")
- print(f"{'PDB Atlas':<25} {results_df['pdb_top1_extended'].mean()*100:>14.1f}% {'+' + str(round(results_df['pdb_top1_extended'].mean()*100 - random_pdb_extended, 1)):>11}%")
- print(f"\nAll saved files:")
- print(" - baseline_comparison_summary.csv")
- print(" - baseline_comparison_cohort.csv")
- print(" - baseline_comparison_detailed.csv")
- print(" - baseline_mcnemar_tests.csv")
- print(" - baseline_confidence_margins.csv")
- print(" - baseline_topk_accuracy_lines.png/svg")
- print(" - baseline_cohort_barplot.png/svg")
- print(" - baseline_cohort_topk_lines.png/svg")
- print(" - baseline_accuracy_heatmap.png/svg")
- print(" - baseline_confusion_matrices.png/svg")
- print(" - baseline_umap_centroids.png/svg")
- print(" - baseline_confidence_margins.png/svg")
- print(" - baseline_final_comparison.png/svg")
- # %%
baseline_comparison_clean.ipynb at commit c164143, no license · at the source
Overview
- VIB-UGent Center for Medical Biotechnology, VIB,Ghent, Belgium
- Department of Biomolecular Medicine, Ghent University,Ghent, Belgium
- BioOrganic Mass Spectrometry Laboratory (LSMBO), UMR 7178, IPHC, University of Strasbourg, CNRS,Strasbourg, 67000 France
- Infrastructure Nationale de Protéomique, ProFI-UAR 2048, Strasbourg, France
Abstract
MLMarker is a machine learning tool that computes continuous tissue similarity scores for proteomics data, addressing the challenge of interpreting complex or sparse datasets. Trained on 34 healthy tissues, its Random Forest model generates probabilistic predictions with SHAP-based protein-level explanations. A penalty factor corrects for missing proteins, improving robustness for low-coverage samples. Across three public datasets, MLMarker revealed brain-like signatures in cerebral melanoma metastases, achieved high accuracy in a pan-cancer cohort, and identified brain and pituitary origins in biofluids. MLMarker provides an interpretable framework for tissue inference and hypothesis generation, available as a Python package and Streamlit app.
Supplementary Information: The online version contains supplementary material available at 10.1186/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 34 matches between paragraphs and lines of code.
TineClaeys/MLMarker-manuscript
c164143cc2859301d5dff8ab50767401fcf72ba8, 13 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
24 files
- Cancer_datasets_Kuster/
baseline_comparison_clea , Jupyter, 2,059 lines, 4 matchesn.ipynb - Cancer_datasets_Kuster/
brain_melanoma.ipynb , Jupyter, 260 lines - Cancer_datasets_Kuster/
tissue_similarity_visual , Jupyter, 372 linesization.ipynb - Figures_paper/
figure_style.py , Python, 275 lines - Figures_paper/
manuscript_style.py , Python, 565 lines - MLMarker_reprocessing_qu
ant.py , Python, 492 lines - Missingnes_testing/
01_setupfiles.ipynb , Jupyter, 678 lines - Missingnes_testing/
02_penalty_factor_compar , Jupyter, 584 lines, 4 matchesison.ipynb - Missingnes_testing/
penalty_benefit_figure.p , Python, 260 linesy - Missingnes_testing/
penalty_factor_analysis. , Python, 1,160 lines, 3 matchespy - Missingnes_testing/
rank_analysis_heatmap.py , Python, 294 lines, 2 matches - PXD007592_responders_mel
anoma/ , Jupyter, 109 lines01_data_preprocessing.ip ynb - PXD007592_responders_mel
anoma/ , Jupyter, 736 lines, 4 matches03_shap_vs_dea_correlati on.ipynb - PXD007592_responders_mel
anoma/ , Jupyter, 471 lines04_figures.ipynb - PXD007592_responders_mel
anoma/ , Jupyter, 376 lines, 2 matches05_clusteringgroups.ipyn b - PXD007592_responders_mel
anoma/ , Jupyter, 950 linesbrain_vs_notbrain_invasi veness.ipynb - PXD007592_responders_mel
anoma/ , Jupyter, 270 linesclassify_01_setup_mlmark er.ipynb - PXD007592_responders_mel
anoma/ , Jupyter, 174 linesclassify_02_umap_cluster ing.ipynb - PXD007592_responders_mel
anoma/ , Jupyter, 216 linesclassify_03_go_enrichmen t.ipynb - PXD007592_responders_mel
anoma/ , Jupyter, 911 lines, 1 matchclassify_04_protein_anal ysis.ipynb - PXD007592_responders_mel
anoma/ , Jupyter, 404 linesclassify_05_literature_c omparison.ipynb - PXD007592_responders_mel
anoma/ , R, 223 lines, 2 matchesmsqrob_analysis.R - PXD007592_responders_mel
anoma/ , R, 422 linesmsqrob_analysis_unified. R - PXD008029_CSF/
classify_csf.ipynb , Jupyter, 711 lines, 2 matches - repository limit reached (2,000 files or 30 MB): the rest is at the source (1 files)
TineClaeys/MLMarker
6444accee481886a09fd04c9b873ae1185c23420, 30 January 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
18 files
- MLMarker_Arnaud_quant.py
, Python, 139 lines - Tissue_predictors2_fulld
ata.ipynb , Jupyter, 1,265 lines - mlmarker/
__init__.py , Python, 6 lines - mlmarker/
constants.py , Python, 15 lines - mlmarker/
explainability.py , Python, 290 lines, 1 match - mlmarker/
model.py , Python, 134 lines - mlmarker/
utils.py , Python, 291 lines - old_modules/
atlas.py , Python, 144 lines - old_modules/
cell_predictors.py , Python, 70 lines - old_modules/
database.py , Python, 164 lines - old_modules/
mlmarker.py , Python, 199 lines, 1 match - old_modules/
predictor_atlas.py , Python, 101 lines - old_modules/
tissue_predictors.py , Python, 69 lines - setup.py, Python, 34 lines
- test_notebook.ipynb, Jupyter, 204 lines
- tests/
test_mlmarker.py , Python, 141 lines - LICENSE, License, 201 lines
- README.md, Text, 145 lines
TineClaeys/MLMarker-streamlit
2bf8a28f94efe81f9d89ddb6223d255b1e3af9e0, 21 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
12 files
- Home.py, Python, 625 lines, 2 matches
- Markverse/
crop_images.py , Python, 33 lines - classify_04_protein_anal
ysis.ipynb , Jupyter, 906 lines - custom_functions.py, Python, 659 lines, 2 matches
- pages/
1_QC.py , Python, 564 lines, 2 matches - pages/
2_Visualisations.py , Python, 246 lines, 1 match - pages/
3_ProteinExplorer.py , Python, 267 lines - pages/
4_FunctionalAnalysis.py , Python, 274 lines - pages/
5_Comparison.py , Python, 886 lines - test.ipynb, Jupyter, 148 lines, 1 match
- LICENSE, License, 201 lines
- README.md, Text, 3 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 50 scripts, each with its path and the digest of its content;
- 34 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- figshare:32787522, at figshare; found in DataCite
Data availability
All datasets used in this study are publicly available via ProteomeXchange (PXD007592, PXD009021, PXD008029) and MassIVE (MSV000095036). MLMarker predictions, differential expression results, and SHAP-derived protein lists are available in the Supplementary Data & Github repository for the manuscript analyses, pip package and streamlit application respectively github.com/
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Authors: added Tine Claeys (0000-0001-9408-488X); Sam van Puyenbroeck (0009-0001-5281-6891); Kris Gevaert (0000-0002-4237-0283); Lennart Martens (0000-0003-4277-658X); removed Tine Claeys; Sam van Puyenbroeck; Kris Gevaert; Lennart Martens
- Funding: added CHIST-ERA: G0GDV23N
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 5 keywords, 8 MeSH terms, 41 references.
Cite
This paper
Claeys, T., van Puyenbroeck, S., Gevaert, K., & Martens, L. (2026). MLMarker: a machine learning framework for tissue inference and biomarker discovery. Genome biology, 27(1), 207. https://
BibTeX
@article{claeys2026mlmar
author = {Claeys, Tine and van Puyenbroeck, Sam and Gevaert, Kris and Martens, Lennart},
title = {{MLMarker: a machine learning framework for tissue inference and biomarker discovery}},
journal = {Genome biology},
year = {2026},
month = jun,
volume = {27},
number = {1},
pages = {207},
publisher = {BMC},
issn = {1474-7596},
doi = {10.1186/
url = {https://
pmid = {42343371},
pmcid = {PMC13292323}
}
RIS
TY - JOUR
AU - Claeys, Tine
AU - van Puyenbroeck, Sam
AU - Gevaert, Kris
AU - Martens, Lennart
TI - MLMarker: a machine learning framework for tissue inference and biomarker discovery
T2 - Genome biology
J2 - Genome Biol
PY - 2026
DA - 2026/
VL - 27
IS - 1
SP - 207
SN - 1474-7596
PB - BMC
DO - 10.1186/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1186/
"type": "article-journal",
"title": "MLMarker: a machine learning framework for tissue inference and biomarker discovery",
"container-title": "Genome biology",
"author": [
{
"family": "Claeys",
"given": "Tine"
},
{
"family": "van Puyenbroeck",
"given": "Sam"
},
{
"family": "Gevaert",
"given": "Kris"
},
{
"family": "Martens",
"given": "Lennart"
}
],
"container-title-short":
"volume": "27",
"issue": "1",
"page": "207",
"DOI": "10.1186/
"PMID": "42343371",
"PMCID": "PMC13292323",
"ISSN": "1474-7596",
"publisher": "BMC",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
24
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: imbalanced-learn, SHAP, XGBoost, 13 other tools, 1 reference
- [2] doi:10.1016/j.isci.2026.116825 [code]
- Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.Journal: iScienceIn common: imbalanced-learn, SHAP, XGBoost, 11 other tools, genetics / omics
- [3] doi:10.7554/elife.110588 [code]
- Opening the black box toward a modular approach to spike sorting.Journal: eLifeIn common: imbalanced-learn, XGBoost, UMAP, 8 other tools, methods / tools
- [4] doi:10.1038/s41467-026-76939-w [code]
- HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities.Journal: Nature communicationsIn common: imbalanced-learn, UMAP, Pillow, 9 other tools, methods / tools
- [5] doi:10.1016/j.isci.2026.115329 [code]
- Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.Journal: iScienceIn common: imbalanced-learn, SHAP, XGBoost, 8 other tools, genetics / omics
- [6] doi:10.1016/j.nicl.2026.104012 [code]
- Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.Journal: NeuroImage. ClinicalIn common: imbalanced-learn, SHAP, XGBoost, 8 other tools
- [7] doi:10.1038/s43856-026-01606-6 [code]
- Validation of remote multimodal AI screening for Parkinson disease across diverse settings.Journal: Communications medicineIn common: imbalanced-learn, SHAP, Plotly, 8 other tools
- [8] doi:10.1126/sciadv.aed3650 [code]
- Truthful visualizations for mass spectrometry imaging enable high-spatial-resolution interactive &
lt;i& gt;m/ z& lt;/ i& gt; mapping and exploration. Journal: Science advancesIn common: SHAP, XGBoost, UMAP, 7 other tools, methods / tools, genetics / omics - [9] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: limma, UMAP, Plotly, 8 other tools, genetics / omics
- [10] doi:10.3390/s26175327 [code]
- Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER.Journal: Sensors (Basel, Switzerland)In common: imbalanced-learn, SHAP, XGBoost, 7 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 3 repositories of the authors' code, each at its verified commit and with its license, 50 scripts, and 34 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:2e990cd5fd04b01b…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
