Machine learning-based combination of the central vein sign, cortical lesions and paramagnetic rim lesions: a web-based tool for the diagnosis of multiple sclerosis.
The 7 matches
- [1] § Materials and methods › Machine learning framework and evaluation protocol ↔ step_1_find_best_algorithm_for_each_combination_of_features.py, lines 64–114 · score 0.85 · neighbours classifier, decision tree, logistic regression, random forest, boosting, balanced accuracy
- [2] § Materials and methods › Machine learning framework and evaluation protocol ↔ interpretability_plots.py, lines 52–76 · score 0.62 · logistic regression, random forest, tree, SVC, XGB, classification
- [3] § Materials and methods › Machine learning framework and evaluation protocol ↔ step_6_retrain_models_on_full_training_set.py, lines 50–109 · score 0.57 · F1 score, full training, retrained, balanced accuracy, precision, sensitivity
- [4] § Results ↔ step_1_find_best_algorithm_for_each_combination_of_features.py, lines 64–114 · score 0.56 · random forest classifier, logistic regression, best model, prediction
- [5] § Results ↔ interpretability_plots.py, lines 207–253 · score 0.55 · random forest classifier, logistic regression, SHAP, prediction, model
- [6] § Materials and methods › Machine learning framework and evaluation protocol ↔ step_2_find_subset_of_algorithm_combination_pairs_outperforming_dis.py, lines 159–247 · score 0.53 · simplified variable, algorithm combination pairs, CL1, PRL1, DIS, balanced accuracy
- [7] § Results ↔ step_2_find_subset_of_algorithm_combination_pairs_outperforming_dis.py, lines 159–247 · score 0.51 · outperformed DIS, algorithm combination pairs, balanced accuracy, simplified, variables, CVS
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 188 lines · 7.4 KB · no license · 2 matches
- import pandas as pd
- import json
- import numpy as np
- from sklearn.model_selection import train_test_split
- from sklearn.preprocessing import StandardScaler
- from sklearn.metrics import accuracy_score, confusion_matrix, classification_report, balanced_accuracy_score, f1_score
- from sklearn.linear_model import LogisticRegression
- from sklearn.svm import SVC
- from sklearn.ensemble import RandomForestClassifier
- from sklearn.model_selection import GridSearchCV
- from sklearn.pipeline import Pipeline
- from sklearn.neighbors import KNeighborsClassifier
- from sklearn.tree import DecisionTreeClassifier
- from xgboost import XGBClassifier
- from sklearn.ensemble import AdaBoostClassifier
- from itertools import chain, combinations
- import warnings
- from utils import *
- warnings.filterwarnings("ignore", category=FutureWarning)
- from sklearn.model_selection import cross_val_predict, StratifiedKFold, KFold
- from sklearn.metrics import balanced_accuracy_score, confusion_matrix
- import numpy as np
- global n_model
- n_model = 0
- np.random.seed(RANDOM_SEED)
- def powerset(iterable):
- "powerset([1,2,3]) --> (1,) (2,) (3,) (1,2) (1,3) (2,3) (1,2,3)"
- s = list(iterable)
- return chain.from_iterable(combinations(s, r) for r in range(1, len(s)+1))
- def find_best_threshold(y_true, y_pred_proba):
- """
- Find the best threshold for the given predictions
- :param y_true: GT labels
- :param y_pred_proba: predicted probabilities
- :return: best threshold
- """
- thresholds = np.linspace(0, 1, 100)
- best_threshold = None
- best_balanced_accuracy = -1
- best_sensitivity = -1
- for threshold in thresholds:
- y_pred = (y_pred_proba[:, 1] >= threshold).astype(int)
- tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()
- sensitivity = tp / (tp + fn)
- specificity = tn / (tn + fp)
- balanced_accuracy = (sensitivity + specificity) / 2
- if balanced_accuracy > best_balanced_accuracy or \
- (balanced_accuracy == best_balanced_accuracy and sensitivity > best_sensitivity):
- best_threshold = threshold
- best_balanced_accuracy = balanced_accuracy
- best_sensitivity = sensitivity
- return best_threshold
- def cross_validate_best_model(X, y):
- models = [
- LogisticRegression(),
- SVC(probability=True),
- RandomForestClassifier(),
- KNeighborsClassifier(),
- DecisionTreeClassifier(),
- XGBClassifier(),
- AdaBoostClassifier(),
- ]
- best_model = None
- best_balanced_accuracy = -1
- associated_probability_threshold = None
- for model_class in models:
- skf = StratifiedKFold(n_splits=10, shuffle=True, random_state=RANDOM_SEED)
- # skf = KFold(n_splits=10, shuffle=True, random_state=42)
- balanced_accuracies = []
- probability_thresholds = []
- for train_index, test_index in skf.split(X, y):
- X_train, X_test = X.iloc[train_index], X.iloc[test_index]
- y_train, y_test = y[X_train.index], y[X_test.index]
- if USE_PIPELINE:
- model = Pipeline([
- ("scaler", StandardScaler()),
- ("model", model_class)
- ])
- else:
- model = model_class
- model.fit(X_train, y_train)
- y_pred_proba = model.predict_proba(X_test)
- best_threshold = find_best_threshold(y_test, y_pred_proba)
- y_pred = (y_pred_proba[:, 1] >= best_threshold).astype(int)
- balanced_accuracies.append(balanced_accuracy_score(y_test, y_pred))
- probability_thresholds.append(best_threshold)
- avg_balanced_accuracy = np.mean(balanced_accuracies)
- if avg_balanced_accuracy > best_balanced_accuracy:
- best_model = model_class
- best_balanced_accuracy = avg_balanced_accuracy
- associated_probability_threshold = np.mean(probability_thresholds)
- best_model = str(best_model).split("(")[0]
- best_model = 'XGBClassifier()' if "XGBClassifier" in str(best_model) else best_model
- return best_model, associated_probability_threshold
- def find_best_algorithm_for_specific_combination_of_features(X, y, features):
- global n_model
- n_model += 1
- feats = [f for f in features if f not in ("Select3*-v2NA", "Select6*-v2NA")]
- if "Select3*-v2NA" in features:
- feats += ["Select3*-v2NA_0.0", "Select3*-v2NA_1.0", "Select3*-v2NA_NA"]
- if "Select6*-v2NA" in features:
- feats += ["Select6*-v2NA_0.0", "Select6*-v2NA_1.0", "Select6*-v2NA_NA"]
- this_X = X[feats].dropna()
- this_y = y[this_X.index]
- # # #reindex
- # # this_X = this_X.reset_index(drop=True)
- # # this_y = this_y.reset_index(drop=True)
- # this_X, this_y = prepare_data(this_X, this_y)
- best_model, best_threshold = cross_validate_best_model(this_X, this_y)
- return best_model, best_threshold
- def make_combinations(excluded_features):
- all_features = [f for f in ALL_FEATURES if f not in excluded_features]
- all_combinations = powerset(all_features)
- n_combinations = 2 ** len(all_features) - 1
- print(f"Number of combinations: {n_combinations}")
- # remove all combinations where both Select3*-v2NA and % perivenular les are present
- all_combinations = [c for c in all_combinations if not ("Select3*-v2NA" in c and "% perivenular les" in c)]
- # remove all combinations where both Select6*-v2NA and % perivenular les are present
- all_combinations = [c for c in all_combinations if not ("Select6*-v2NA" in c and "% perivenular les" in c)]
- # remove all combinations where both PRL1 and number_PRL are present
- all_combinations = [c for c in all_combinations if not ("PRL1" in c and "number_PRL" in c)]
- # remove all combinations where both CL1 and CL-count-updated are present
- all_combinations = [c for c in all_combinations if not ("CL1" in c and "CL-count-updated" in c)]
- # remove all combinations where both Select3*-v2NA and Select6*-v2NA are present
- all_combinations = [c for c in all_combinations if not ("Select3*-v2NA" in c and "Select6*-v2NA" in c)]
- print(f"Number of combinations after filtering: {len(all_combinations)}")
- return all_combinations
- def find_best_algorithm_for_each_combination(X, y):
- excluded_features = ("age", "sex", "Filippi", "OCB_presence")
- all_features = [f for f in ALL_FEATURES if f not in excluded_features]
- all_combinations = make_combinations(excluded_features)
- best_model_per_combination = []
- for i, combination in enumerate(all_combinations):
- print(f"Progress: {i + 1}/{len(all_combinations)}")
- best_model, best_threshold = find_best_algorithm_for_specific_combination_of_features(X, y, combination)
- print(f"Best model for combination {combination}: {best_model}, best threshold: {best_threshold}\n")
- feature_presence = {f: f in list(combination) for f in sorted(all_features, key=lambda x: x.upper())}
- infos = {
- "model": best_model,
- "threshold": best_threshold,
- }
- infos.update(feature_presence)
- best_model_per_combination.append(infos)
- best_model_per_combination = pd.DataFrame(best_model_per_combination)
- return best_model_per_combination
- if __name__ == "__main__":
- X, y = get_full_dataset(TRAINING_DATA_PATH, dropna=False, return_X_y=True, prepare=True)
- best_models = find_best_algorithm_for_each_combination(X, y)
- best_models.to_csv(BEST_ALGORITHM_PER_COMBINATION_OF_FEATURES_PATH, index=False)
- print("Done.")
step_1_find_best_algorithm_for_each_combination_of_features.py at commit 9551af5, no license · at the source
Overview
14 affiliations
- ICTEAM Institute, Université Catholique de Louvain, 1348 Louvain-la-Neuve, Belgium
- Neuroinflammation Imaging Lab (NIL), Université Catholique de Louvain, 1200 Brussels, Belgium
- Department of Neurology, Cliniques Universitaires Saint-Luc (CUSL), Université Catholique de Louvain, 1200 Brussels, Belgium
- Department of Neurology, Hôpital Erasme, Hôpital Universitaire de Bruxelles, Université Libre de Bruxelles, 1070 Brussels, Belgium
- CIBM Center for Biomedical Imaging, CH-1015 Lausanne, Switzerland
- Radiology Department, Lausanne University Hospital (CHUV) and University of Lausanne, CH-1011 Lausanne, Switzerland
- Neurology Unit, IRCCS San Raffaele Hospital, 20132 Milan, Italy
- Department of Neurosciences and Biomedicine and Movement, The Multiple Sclerosis Center of University Hospital of Verona, 37129 Verona, Italy
- Department of Neurology, Cedars-Sinai Medical Center, 90048 Los Angeles, CA, USA
- Vita-Salute San Raffaele University, Milan, Italy
- Neuroimaging Research Unit, Division of Neuroscience, IRCCS San Raffaele Scientific Institute, 20132 Milan, Italy
- Experimental Neuropathology Lab, IRCCS Humanitas Research Institute, 20132 Milan, Italy
- Department of Biomedical Sciences, Humanitas University, 20072 Milan, Italy
- Translational Neuroradiology Section, National Institute of Neurological Disorders and Stroke (NINDS), National Institutes of Health (NIH), 20892 Bethesda, MD, USA
Abstract
Multiple sclerosis diagnostic criteria lack optimal specificity, leading to potential misdiagnosis. Advanced magnetic resonance imaging (MRI) biomarkers like the central vein sign, cortical lesions and paramagnetic rim lesions are highly specific to multiple sclerosis and could potentially improve diagnostic accuracy. In this study, we applied machine learning techniques to a retrospective, multicentric dataset of 322 multiple sclerosis/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.
maxencewynen/MS-Diagnostic-Tool-UCLouvain
9551af5d43f011334c8dd8734b34c3424022cf73, 25 September 2024Availability: 1 check, the latest on 30 September 2026: the link answers
- 30 September 2026: the link answers
9 files
- interpretability_plots.p
y , Python, 263 lines, 2 matches - step_1_find_best_algorit
hm_for_each_combination_ , Python, 188 lines, 2 matchesof_features.py - step_2_find_subset_of_al
gorithm_combination_pair , Python, 277 lines, 2 matchess_outperforming_dis.py - step_3_find_best_algorit
hm_combination_pair.py , Python, 16 lines - step_4_find_best_algorit
hm_combination_pair_simp , Python, 27 lineslified_variables.py - step_5_compare_best_over
all_model_vs_best_simpli , Python, 19 linesfied_model.py - step_6_retrain_models_on
_full_training_set.py , Python, 114 lines, 1 match - step_7_infer_on_test_set
_and_compute_metrics.py , Python, 130 lines - utils.py, Python, 426 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 9 scripts, each with its path and the digest of its content;
- 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
Qualified researchers can access the individual patient data of this study upon reasonable request and material transfer agreement between institutes. The trained models of the study are available at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 30 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 17 authors, 5 keywords, 19 funders, 25 references.
Cite
This paper
Wynen, M., Vanden Bulcke, C., Borrelli, S., Gordaliza, P. M., Stölting, A., Guisset, F., Cordier, C., Martire, M. S., Tamanti, A., Macq, B., Sati, P., Filippi, M., Calabrese, M., Absinta, M., Reich, D. S., Bach Cuadra, M., & Maggi, P. (2026). Machine learning-based combination of the central vein sign, cortical lesions and paramagnetic rim lesions: a web-based tool for the diagnosis of multiple sclerosis. Brain communications, 8(2), fcag079. https://
BibTeX
@article{wynen2026machin
author = {Wynen, Maxence and Vanden Bulcke, Colin and Borrelli, Serena and Gordaliza, Pedro M and Stölting, Anna and Guisset, François and Cordier, Clément and Martire, Maria Sofia and Tamanti, Agnese and Macq, Benoit and Sati, Pascal and Filippi, Massimo and Calabrese, Massimiliano and Absinta, Martina and Reich, Daniel S and Bach Cuadra, Meritxell and Maggi, Pietro},
title = {{Machine learning-based combination of the central vein sign, cortical lesions and paramagnetic rim lesions: a web-based tool for the diagnosis of multiple sclerosis}},
journal = {Brain communications},
year = {2026},
month = mar,
volume = {8},
number = {2},
pages = {fcag079},
publisher = {Oxford University Press},
issn = {2632-1297},
doi = {10.1093/
url = {https://
pmid = {41884595},
pmcid = {PMC13010066}
}
RIS
TY - JOUR
AU - Wynen, Maxence
AU - Vanden Bulcke, Colin
AU - Borrelli, Serena
AU - Gordaliza, Pedro M
AU - Stölting, Anna
AU - Guisset, François
AU - Cordier, Clément
AU - Martire, Maria Sofia
AU - Tamanti, Agnese
AU - Macq, Benoit
AU - Sati, Pascal
AU - Filippi, Massimo
AU - Calabrese, Massimiliano
AU - Absinta, Martina
AU - Reich, Daniel S
AU - Bach Cuadra, Meritxell
AU - Maggi, Pietro
TI - Machine learning-based combination of the central vein sign, cortical lesions and paramagnetic rim lesions: a web-based tool for the diagnosis of multiple sclerosis
T2 - Brain communications
J2 - Brain Commun
PY - 2026
DA - 2026/
VL - 8
IS - 2
SP - fcag079
SN - 2632-1297
PB - Oxford University Press
DO - 10.1093/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1093/
"type": "article-journal",
"title": "Machine learning-based combination of the central vein sign, cortical lesions and paramagnetic rim lesions: a web-based tool for the diagnosis of multiple sclerosis",
"container-title": "Brain communications",
"author": [
{
"family": "Wynen",
"given": "Maxence"
},
{
"family": "Vanden Bulcke",
"given": "Colin"
},
{
"family": "Borrelli",
"given": "Serena"
},
{
"family": "Gordaliza",
"given": "Pedro M"
},
{
"family": "Stölting",
"given": "Anna"
},
{
"family": "Guisset",
"given": "François"
},
{
"family": "Cordier",
"given": "Clément"
},
{
"family": "Martire",
"given": "Maria Sofia"
},
{
"family": "Tamanti",
"given": "Agnese"
},
{
"family": "Macq",
"given": "Benoit"
},
{
"family": "Sati",
"given": "Pascal"
},
{
"family": "Filippi",
"given": "Massimo"
},
{
"family": "Calabrese",
"given": "Massimiliano"
},
{
"family": "Absinta",
"given": "Martina"
},
{
"family": "Reich",
"given": "Daniel S"
},
{
"family": "Bach Cuadra",
"given": "Meritxell"
},
{
"family": "Maggi",
"given": "Pietro"
}
],
"container-title-short":
"volume": "8",
"issue": "2",
"page": "fcag079",
"DOI": "10.1093/
"PMID": "41884595",
"PMCID": "PMC13010066",
"ISSN": "2632-1297",
"publisher": "Oxford University Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
11
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.nicl.2026.104007 [code]
- A comparative study of deep learning for cortical lesion MRI segmentation with explainability analysis in multiple sclerosis.Journal: NeuroImage. ClinicalIn common: multiple sclerosis, structural MRI / diffusion, 4 references, 2 authors
- [2] doi:10.64898/2026.07.15.26357954 [code]
- Portable Ultra-Low Field MRI Deep-Learning Algorithms for White Matter Lesion Segmentation Improve Accuracy and Reflect Clinical Disability in Multiple SclerosisJournal: medRxiv (preprint)In common: seaborn, scikit-learn, pandas, 3 other tools, multiple sclerosis, structural MRI / diffusion, clinical / translational, 1 reference, author Daniel S. Reich
- [3] doi:10.1136/jnnp-2025-335884 [code]
- Diffusivity anisotropy signature of slowly expanding lesions predicts progression independent of relapse activity in multiple sclerosis.Journal: Journal of neurology, neurosurgery, and psychiatryIn common: seaborn, scikit-learn, pandas, 3 other tools, multiple sclerosis, structural MRI / diffusion, clinical / translational, 3 references
- [4] doi:10.1186/s12880-026-02481-2 [code]
- Deep learning-based neuroanatomical profiling reveals population-specific brain changes in multiple sclerosis: a large-scale Middle Eastern study.Journal: BMC medical imagingIn common: seaborn, scikit-learn, pandas, 3 other tools, multiple sclerosis, structural MRI / diffusion, clinical / translational, 3 references
- [5] doi:10.7554/elife.109717 [code]
- Retrosplenial cortex enables context-dependent goal-directed sensorimotor transformation.Journal: eLifeIn common: SHAP, XGBoost, seaborn, 5 other tools, 1 reference
- [6] doi:10.1073/pnas.2516601123 [code]
- Unveiling the glymphatic system's role in brain aging: A comprehensive biomarker and modifiable intervention target.Journal: Proceedings of the National Academy of Sciences of the United States of AmericaIn common: SHAP, XGBoost, seaborn, 5 other tools, structural MRI / diffusion, clinical / translational
- [7] doi:10.1038/s41467-026-71555-0 [code]
- A deep representation learning model to predict response to vagus nerve stimulation.Journal: Nature communicationsIn common: SHAP, XGBoost, seaborn, 5 other tools, structural MRI / diffusion, clinical / translational
- [8] doi:10.21203/rs.3.rs-9914920/v1 [code]
- Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working GroupJournal: Research Square (preprint)In common: SHAP, XGBoost, seaborn, 5 other tools, structural MRI / diffusion
- [9] doi:10.1016/j.nicl.2026.104012 [code]
- Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.Journal: NeuroImage. ClinicalIn common: SHAP, XGBoost, seaborn, 5 other tools, structural MRI / diffusion
- [10] doi:10.1016/j.isci.2026.115329 [code]
- Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.Journal: iScienceIn common: SHAP, XGBoost, seaborn, 5 other tools, clinical / translational
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 9 scripts, and 7 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:85ccb43855944b28…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
