Feature selection leads to divergent neurobiological interpretations of brain-based machine learning biomarkers.
The 2 matches
- [1] § Methods › Ridge regression CPM ↔ train_model_ranked_edges.m, lines 221–259 · score 0.55 · hyperparameter optimization, fitrlinear, fitting, training, ridge, CPM
- [2] § Methods › Datasets ↔ run_models.m, lines 38–46 · score 0.51 · Card Sort age, Flanker, NIH, HBN
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
MATLAB · 499 lines · 19 KB · CC-BY-4.0 · 1 match
- function results = train_model(dataset, varargin)
- %{
- Function to train a ridge CPM or basic CPM model
- To use this script for within dataset predictions, you should pass one
- dataset
- To use this script for external (cross-dataset) predictions, you should
- pass two datasets
- %}
- %% Parse input arguments
- p = inputParser;
- addRequired(p, 'dataset', @isstruct);
- addParameter(p, 'external_dataset', '', @isstruct); % external pca table
- addParameter(p, 'model_type', 'ridge', @ischar); % options: ridge, cpm
- addParameter(p, 'seed', 1, @isnumeric);
- addParameter(p, 'num_folds', 10, @isnumeric);
- addParameter(p, 'feat_thresh', 0.05, @isnumeric);
- addParameter(p, 'feat_selection', 'percent', @ischar); % percent, p, ranked_10_percent, ranked_20_percent, ranked_5_percent, ranked_1_percent, all, top_half, bottom_half
- addParameter(p, 'null', 0, @isnumeric); % 1 for yes
- addParameter(p, 'control_covars', 0, @isnumeric); % 1 for yes
- addParameter(p, 'ranked_segment', 1, @isnumeric);
- addParameter(p, 'lambda', NaN, @isnumeric);
- parse(p, dataset, varargin{:});
- external_dataset = p.Results.external_dataset;
- model_type = p.Results.model_type;
- seed = p.Results.seed;
- num_folds = p.Results.num_folds;
- feat_thresh = p.Results.feat_thresh;
- feat_selection = p.Results.feat_selection;
- null = p.Results.null;
- control_covars = p.Results.control_covars == 1;
- ranked_segment = p.Results.ranked_segment;
- lambda = p.Results.lambda;
- if ~isnan(lambda) && length(lambda) > 1
- error(['Only a single lambda value is supported. ', ...
- 'To optimize lambda automatically, omit the lambda parameter or set it to NaN.']);
- end
- clearvars p
- %% get data
- % matrix data (vectorized)
- X = mat2edge(dataset.mats)';
- % behavioral data
- num_behav = length(dataset.phenotypes);
- behav_all = dataset.behav_table{:, dataset.phenotypes};
- n = length(behav_all);
- % covariate data
- if control_covars
- if isfield(dataset, 'covars') && isfield(dataset, 'covar_table')
- covars = dataset.covar_table{:, dataset.covars};
- disp('running with covariates')
- else
- disp('control_covars is set to run covariates, but either covars phenotypes or covar_table not specified')
- end
- end
- % shuffle for null distribution
- rng(seed)
- if null == 1
- shuffle_idx = randperm(length(behav_all));
- behav_all = behav_all(shuffle_idx, :);
- if control_covars && isfield(dataset, 'covars') && isfield(dataset, 'covar_table')
- covars = covars(shuffle_idx, :);
- end
- end
- % option for no PCA, external PCA, or cross-validated PCA
- if length(dataset.phenotypes) == 1
- pca_type = 'none';
- behav = behav_all;
- elseif (length(dataset.phenotypes) > 1) && (isfield(dataset, 'external_behav_table'))
- % if more than 1 phenotype present and if external pca file exists
- pca_type = 'external';
- % run pca in external (non-imaging) data
- behav_external_pca = dataset.external_behav_table{:, dataset.phenotypes};
- [pca_coeff, score, latent, tsquared, explained, mu] = pca(behav_external_pca);
- % project behavior in main dataset
- behav_pcs = (behav_all - mu) * pca_coeff;
- behav = behav_pcs(:, 1);
- elseif (length(dataset.phenotypes) > 1) && (~isfield(dataset, 'external_behav_table'))
- % if more than 1 phenotype present but no external pca file
- if ~isstruct(external_dataset)
- % do PCA within cross-validation folds for within-dataset predictions
- pca_type = 'cv';
- % make an empty array to fill in with PC test data (obtained via PCA within cross-validation scheme)
- behav = NaN + zeros(n, 1);
- elseif isstruct(external_dataset)
- % if no behavior-only pca file is provided, you can still do internal PCA when predicting in an external dataset
- pca_type = 'internal';
- % run pca in external (non-imaging) data
- [pca_coeff, score, latent, tsquared, explained, mu] = pca(behav_all);
- % project behavior in main dataset
- behav_pcs = (behav_all - mu) * pca_coeff;
- behav = behav_pcs(:, 1);
- end
- end
- rng(seed)
- %% Within-dataset
- if isempty(external_dataset)
- % Select cross-validation splits
- cv_idx = cv_indices(n, num_folds);
- coef_all = zeros(size(X, 2), num_folds);
- coef0_all = zeros(num_folds, 1);
- y_predict = zeros(n, 1);
- lambda_total = NaN(num_folds, 1);
- for k = 1:num_folds
- train_idx = find(cv_idx ~= k);
- test_idx = find(cv_idx == k);
- X_train = X(train_idx, :);
- X_test = X(test_idx, :);
- if strcmp(pca_type, 'cv')
- % run pca in training data
- [pca_coeff, score, latent, tsquared, explained, mu] = pca(behav_all(train_idx, :));
- % project behavior in training data
- behav_pcs = (behav_all(train_idx, :) - mu) * pca_coeff;
- behav_train = behav_pcs(:, 1);
- % project behavior in test data
- behav_pcs = (behav_all(test_idx, :) - mu) * pca_coeff;
- behav_test = behav_pcs(:, 1);
- behav(test_idx) = behav_test; % save the test cross-validated PC for later
- else
- behav_train = behav(train_idx);
- behav_test = behav(test_idx);
- end
- % Step 1 of feature selection: correlation
- if control_covars
- covars_train = covars(train_idx);
- % select features with partial correlation
- [edge_corr, edge_p] = partialcorr(X_train, behav_train, covars_train);
- else
- % select features with standard correlation
- [edge_corr, edge_p] = corr(X_train, behav_train);
- end
- % Step 2 of feature selection: thresholding
- if strcmp(feat_selection, 'p') % selecting by p value
- p_thresh = feat_thresh;
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'percent')
- p_thresh = prctile(edge_p, 100 * feat_thresh);
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'ranked_10_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.10);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'ranked_20_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.20);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'ranked_5_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.05);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'ranked_1_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.01);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'all')
- feat_loc = 1:length(edge_p);
- elseif strcmp(feat_selection, 'top_half')
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(1:round(length(edge_p)/2));
- elseif strcmp(feat_selection, 'bottom_half')
- [~, sorted_indices] = sort(edge_p, 'descend');
- feat_loc = sorted_indices(1:round(length(edge_p)/2));
- elseif strcmp(feat_selection, 'top_10_percent')
- p_thresh = prctile(edge_p, 10);
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'bottom_90_percent')
- p_thresh = prctile(edge_p, 10);
- feat_loc = find(edge_p > p_thresh);
- elseif strcmp(feat_selection, 'top_20_percent')
- p_thresh = prctile(edge_p, 20);
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'bottom_80_percent')
- p_thresh = prctile(edge_p, 20);
- feat_loc = find(edge_p > p_thresh);
- end
- % model fitting
- if strcmp(model_type, 'ridge')
- if isnan(lambda)
- mdl = fitrlinear(X_train(:, feat_loc), behav_train, ...
- 'Learner', 'leastsquares', ...
- 'Regularization', 'ridge', ...
- 'OptimizeHyperparameters', {'Lambda'}, ...
- 'HyperparameterOptimizationOptions', struct('ShowPlots', false, 'Verbose', 0));
- coef = mdl.Beta;
- coef0 = mdl.Bias;
- lambda_total(k) = mdl.ModelParameters.Lambda;
- fprintf('Lambda fold %d: %.4f\n', k, lambda_total(k));
- else
- mdl = fitrlinear(X_train(:, feat_loc), behav_train, ...
- 'Learner', 'leastsquares', ...
- 'Regularization', 'ridge', ...
- 'Lambda', lambda);
- coef = mdl.Beta;
- coef0 = mdl.Bias;
- lambda_total(k) = mdl.Lambda;
- fprintf('Lambda fold %d: %.4f\n', k, lambda_total(k));
- end
- coef_all(feat_loc, k) = coef;
- coef0_all(k) = coef0;
- % prediction
- y_predict(test_idx) = X_test(:, feat_loc) * coef + coef0;
- elseif strcmp(model_type, 'cpm')
- % get positive and negative networks, and summarize into a single feature
- pos_network = feat_loc(edge_corr(feat_loc) > 0);
- neg_network = feat_loc(edge_corr(feat_loc) < 0);
- X_train_summary = sum(X_train(:, pos_network), 2) - sum(X_train(:, neg_network), 2);
- % fit model
- mdl = robustfit(X_train_summary, behav_train);
- % predict
- X_test_summary = sum(X_test(:, pos_network), 2) - sum(X_test(:, neg_network), 2);
- y_predict(test_idx) = mdl(1) + mdl(2) * X_test_summary;
- % save coefficients
- coef = zeros(length(edge_corr), 1);
- coef(pos_network) = mdl(2);
- coef(neg_network) = -mdl(2);
- coef0 = mdl(1);
- coef_all(:, k) = coef;
- coef0_all(k) = coef0;
- end
- end
- results.y_true = behav; % I changed from y to y_true
- results.y_predict = y_predict;
- results.cv_idx = cv_idx;
- coef_consensus_mean = zeros(35778, 1); % Initialize coef_consensus_mean with zeros
- for i = 1:size(coef_all, 1)
- if all(coef_all(i, :) ~= 0)
- coef_consensus_mean(i) = mean(coef_all(i, :));
- end
- end
- results.lambda = lambda_total;
- results.coef_all = coef_all;
- results.coef_consensus_mean = coef_consensus_mean;
- results.coef_mean = mean(coef_all, 2);
- results.coef0_mean = mean(coef0_all, 2);
- results.r = corr(results.y_true, results.y_predict);
- results.pca_type = pca_type;
- results.phenotypes = dataset.phenotypes;
- results.model_type = model_type;
- results.seed = seed;
- results.num_folds = num_folds;
- results.feat_thresh = feat_thresh;
- results.feat_selection = feat_selection;
- results.null = null; % 1 is for null
- results.control_covars = control_covars; % 1 is for yes
- if ismember(feat_selection, {'ranked_10_percent', 'ranked_20_percent', 'ranked_5_percent', 'ranked_1_percent'})
- results.ranked_segment = ranked_segment;
- else
- results.ranked_segment = 'none';
- end
- if control_covars
- results.covars = dataset.covars; % which variables were used as covariates
- else
- results.covars = 'none';
- end
- %% Across datasets
- else
- % get second (external) dataset X data
- X_external = mat2edge(external_dataset.mats)';
- behav_all_external = external_dataset.behav_table{:, external_dataset.phenotypes};
- % for second (external) y data, depends if PCA is needed
- if length(external_dataset.phenotypes)==1
- behav_external = behav_all_external;
- elseif (length(external_dataset.phenotypes)>1) && (isfield(external_dataset, 'external_behav_table'))
- % if more than 1 phenotype present and if external pca file exists
- pca_type = 'external';
- % run pca in external (non-imaging) data
- behav_external_external_pca = external_dataset.external_behav_table{:, external_dataset.phenotypes}; % added "external_" before both "datasets:
- [pca_coeff,score,latent,tsquared,explained,mu] = pca(behav_external_external_pca);
- % project behavior in main dataset
- behav_pcs_external = (behav_all_external-mu)*pca_coeff;
- behav_external = behav_pcs_external(:, 1);
- elseif (length(external_dataset.phenotypes)>1) && (~isfield(external_dataset, 'external_behav_table'))
- pca_type = 'internal';
- % run pca in external dataset
- behav_external_internal_pca = external_dataset.behav_table{:, external_dataset.phenotypes};
- [pca_coeff,score,latent,tsquared,explained,mu] = pca(behav_external_internal_pca);
- % project behavior in main dataset
- behav_pcs_external = (behav_external_internal_pca-mu)*pca_coeff;
- behav_external = behav_pcs_external(:, 1);
- end
- % Step 1 of feature selection: correlation
- if control_covars
- % select features with partial correlation
- [edge_corr, edge_p] = partialcorr(X, behav, covars);
- else
- % select features with standard correlation
- [edge_corr, edge_p] = corr(X, behav);
- end
- % Step 2 of feature selection: thresholding
- if strcmp(feat_selection, 'p')
- p_thresh = feat_thresh;
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'percent')
- p_thresh = prctile(edge_p, 100 * feat_thresh);
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'ranked_10_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.10);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'ranked_20_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.20);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'ranked_5_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.05);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'ranked_1_percent')
- total_features = length(edge_p);
- segment_size = round(total_features * 0.01);
- start_index = (ranked_segment - 1) * segment_size + 1;
- end_index = min(ranked_segment * segment_size, total_features);
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(start_index:end_index);
- elseif strcmp(feat_selection, 'all')
- feat_loc = 1:length(edge_p);
- elseif strcmp(feat_selection, 'top_half')
- [~, sorted_indices] = sort(edge_p);
- feat_loc = sorted_indices(1:round(length(edge_p)/2));
- elseif strcmp(feat_selection, 'bottom_half')
- [~, sorted_indices] = sort(edge_p, 'descend');
- feat_loc = sorted_indices(1:round(length(edge_p)/2));
- elseif strcmp(feat_selection, 'top_10_percent')
- p_thresh = prctile(edge_p, 10);
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'bottom_90_percent')
- p_thresh = prctile(edge_p, 10);
- feat_loc = find(edge_p > p_thresh);
- elseif strcmp(feat_selection, 'top_20_percent')
- p_thresh = prctile(edge_p, 20);
- feat_loc = find(edge_p < p_thresh);
- elseif strcmp(feat_selection, 'bottom_80_percent')
- p_thresh = prctile(edge_p, 20);
- feat_loc = find(edge_p > p_thresh);
- end
- % model training
- if strcmp(model_type, 'ridge')
- if isnan(lambda)
- mdl = fitrlinear(X_train(:, feat_loc), behav_train, ...
- 'Learner', 'leastsquares', ...
- 'Regularization', 'ridge', ...
- 'OptimizeHyperparameters', {'Lambda'}, ...
- 'HyperparameterOptimizationOptions', struct('ShowPlots', false, 'Verbose', 0));
- coef = mdl.Beta;
- coef0 = mdl.Bias;
- lambda_used = mdl.ModelParameters.Lambda;
- fprintf('Lambda used (optimized): %.4f\n', lambda_used);
- else
- mdl = fitrlinear(X(:, feat_loc), behav, ...
- 'Learner', 'leastsquares', ...
- 'Regularization', 'ridge', ...
- 'Lambda', lambda);
- coef = mdl.Beta;
- coef0 = mdl.Bias;
- lambda_used = mdl.Lambda;
- fprintf('Lambda used (fixed): %.4f\n', lambda_used);
- end
- elseif strcmp(model_type, 'cpm')
- pos_network = feat_loc(edge_corr(feat_loc) > 0);
- neg_network = feat_loc(edge_corr(feat_loc) < 0);
- X_summary = sum(X(:, pos_network), 2) - sum(X(:, neg_network), 2);
- % fit model
- mdl = robustfit(X_summary, behav);
- % save coefficients
- coef = zeros(length(edge_corr), 1);
- coef(pos_network) = mdl(2);
- coef(neg_network) = -mdl(2);
- coef0 = mdl(1);
- end
- % prediction
- if strcmp(model_type, 'ridge')
- y_predict = X_external(:, feat_loc) * coef + coef0;
- elseif strcmp(model_type, 'cpm')
- y_predict = X_external * coef + coef0;
- end
- % store results
- results.y_train = behav;
- results.y_external_true = behav_external;
- results.y_external_predict = y_predict;
- results.coef = coef;
- results.coef0 = coef0;
- if strcmp(model_type, 'ridge')
- results.lambda = lambda_used;
- elseif strcmp(model_type, 'cpm')
- results.lambda = NaN;
- end
- results.coef = coef;
- results.coef0 = coef0;
- results.r = corr(results.y_external_true, results.y_external_predict);
- results.pca_type = pca_type;
- results.phenotypes = dataset.phenotypes;
- results.model_type = model_type;
- results.seed = seed;
- results.feat_thresh = feat_thresh;
- results.feat_selection = feat_selection;
- results.null = null; % 1 is for null
- results.control_covars = control_covars; % 1 is for yes
- if ismember(feat_selection, {'ranked_10_percent', 'ranked_20_percent', 'ranked_5_percent', 'ranked_1_percent'})
- results.ranked_segment = ranked_segment; % save the rankem segment number if used
- else
- results.ranked_segment = 'none';
- end
- if control_covars
- results.covars = dataset.covars; % which variables were used as covariates
- else
- results.covars = 'none';
- end
- end
- end
train_model_ranked_edges.m at commit 9b10607, under CC-BY-4.0 · at the source
Overview
- Yale School of Medicine,New Haven, CT USA
- Department of Biomedical Engineering, Yale University,New Haven, CT USA
- Department of Radiology, Perelman School of Medicine,Philadelphia, PA USA
- Department of Psychiatry, University of Pennsylvania,Philadelphia, PA USA
- Department of Radiology & Biomedical Imaging, Yale School of Medicine,New Haven, CT USA
- Department of Psychiatry, University of Oxford, Warneford Hospital,Oxford, UK
- Department of Bioengineering, Northeastern University,Boston, MA USA
- Department of Psychology, Northeastern University,Boston, MA USA
- Institute for Cognitive & Behavioral Health, Northeastern University,Boston, MA USA
- Department of Statistics & Data Science, Yale University,New Haven, CT USA
- Child Study Center, Yale School of Medicine,New Haven, CT USA
- Wu Tsai Institute, Yale University,New Haven, CT USA
Abstract
A central objective in human neuroimaging is to understand the neurobiology underlying cognition and mental health. Machine learning models trained on neuroimaging data are increasingly used as tools for predicting behavioural phenotypes, enhancing precision medicine and improving generalizability compared with traditional MRI studies. However, the high dimensionality of brain connectivity data makes model interpretation challenging. Prevailing practices rely on selecting features and, implicitly, interpreting identified feature networks as uniquely representative of a given phenotype while overlooking others. Despite its widespread use, how univariate feature selection balances the trade-off between simplification for optimizing modelling and oversimplification that misrepresents true neurobiology remains understudied. Here, using four large-scale neuroimaging datasets spanning over 12,000 participants and 13 outcomes, we demonstrate that edges discarded by feature selection can achieve significant prediction accuracies while yielding different neurobiological interpretations. These results are observed across cognitive, developmental and psychiatric phenotypes, extend to both functional connectivity (functional MRI) and structural (diffusion tensor imaging) connectomes, and remain evident in external validation. They suggest that focusing on only the top features may simplify the neurobiological bases of brain–behaviour associations. Such interpretations present only the tip of the iceberg when certain disregarded features may be just as meaningful, potentially contributing to ongoing issues surrounding reproducibility within the field. More broadly, our results reinforce that subtle brain-wide signals should not be ignored.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.
brendan-adkinson/overlooked-features
9b1060783d708e71ee5b5714dd2a8d3a73baf139, 26 September 2025Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
5 files
- cv_indices.m, MATLAB, 16 lines
- mat2edge.m, MATLAB, 11 lines
- run_models.m, MATLAB, 112 lines, 1 match
- train_model_ranked_edges
.m , MATLAB, 499 lines, 1 match - README.md, Text, 62 lines
Code availability
The analyses were conducted using MATLAB v.R2024a. The code is available via GitHub at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 4 scripts, each with its path and the digest of its content;
- 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
The data are available through the PNC, HCPD, HBN and ABCD datasets.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 4 keywords, 10 MeSH terms, 5 funders, 122 references.
Cite
This paper
Adkinson, B. D., Rosenblatt, M., Sun, H., Dadashkarimi, J., Tejavibulya, L., Horien, C., Westwater, M. L., Rodriguez, R. X., Noble, S., & Scheinost, D. (2026). Feature selection leads to divergent neurobiological interpretations of brain-based machine learning biomarkers. Nature human behaviour, 10(7), 1356-1370. https://
BibTeX
@article{adkinson2026fea
author = {Adkinson, Brendan D. and Rosenblatt, Matthew and Sun, Huili and Dadashkarimi, Javid and Tejavibulya, Link and Horien, Corey and Westwater, Margaret L. and Rodriguez, Raimundo X. and Noble, Stephanie and Scheinost, Dustin},
title = {{Feature selection leads to divergent neurobiological interpretations of brain-based machine learning biomarkers}},
journal = {Nature human behaviour},
year = {2026},
month = apr,
volume = {10},
number = {7},
pages = {1356--1370},
publisher = {Nature Portfolio},
issn = {2397-3374},
doi = {10.1038/
url = {https://
pmid = {41986741},
pmcid = {PMC13388108}
}
RIS
TY - JOUR
AU - Adkinson, Brendan D.
AU - Rosenblatt, Matthew
AU - Sun, Huili
AU - Dadashkarimi, Javid
AU - Tejavibulya, Link
AU - Horien, Corey
AU - Westwater, Margaret L.
AU - Rodriguez, Raimundo X.
AU - Noble, Stephanie
AU - Scheinost, Dustin
TI - Feature selection leads to divergent neurobiological interpretations of brain-based machine learning biomarkers
T2 - Nature human behaviour
J2 - Nat Hum Behav
PY - 2026
DA - 2026/
VL - 10
IS - 7
SP - 1356
EP - 1370
SN - 2397-3374
PB - Nature Portfolio
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Feature selection leads to divergent neurobiological interpretations of brain-based machine learning biomarkers",
"container-title": "Nature human behaviour",
"author": [
{
"family": "Adkinson",
"given": "Brendan D."
},
{
"family": "Rosenblatt",
"given": "Matthew"
},
{
"family": "Sun",
"given": "Huili"
},
{
"family": "Dadashkarimi",
"given": "Javid"
},
{
"family": "Tejavibulya",
"given": "Link"
},
{
"family": "Horien",
"given": "Corey"
},
{
"family": "Westwater",
"given": "Margaret L."
},
{
"family": "Rodriguez",
"given": "Raimundo X."
},
{
"family": "Noble",
"given": "Stephanie"
},
{
"family": "Scheinost",
"given": "Dustin"
}
],
"container-title-short":
"volume": "10",
"issue": "7",
"page": "1356-1370",
"DOI": "10.1038/
"PMID": "41986741",
"PMCID": "PMC13388108",
"ISSN": "2397-3374",
"publisher": "Nature Portfolio",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
15
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s44220-026-00623-7 [code]
- Optimizing functional connectivity scanning conditions for predicting autistic traits.Journal: Nature. Mental healthIn common: Statistics and Machine Learning Toolbox, clinical / translational, 12 references, 3 authors
- [2] doi:10.1038/s41467-026-73941-0 [code]
- Using connectome-based predictive models to reveal the systems standardized tests and clinical symptoms are reflecting.Journal: Nature communicationsIn common: 10 references, author Dustin Scheinost
- [3] doi:10.1038/s41467-026-75661-x [code]
- A neural signature of sleep deprivation in the human brain.Journal: Nature communicationsIn common: 12 references
- [4] doi:10.1002/hbm.70493 [code]
- Testing for Network Specificity in Brain-Behavior Associations Using Ordinal Dominance Curves.Journal: Human brain mappingIn common: 8 references
- [5] doi:10.1038/s41467-026-73668-y [code]
- Convergent and divergent brain-cognition development in early adolescence.Journal: Nature communicationsIn common: Statistics and Machine Learning Toolbox, 7 references
- [6] doi:10.7554/elife.104053 [code]
- Brain cognition gaps reveal associations with dopamine and factors related to brain health through artificial intelligence prediction of functional connectome.Journal: eLifeIn common: 7 references
- [7] doi:10.1162/imag.a.1287 [code]
- Is the whole more than the sum of its parts? Considering global and local features of the connectome improves prediction of individuals and phenotypes.Journal: Imaging neuroscience (Cambridge, Mass.)In common: Statistics and Machine Learning Toolbox, 6 references
- [8] doi:10.1038/s41593-026-02359-0 [code]
- The cross-site reproducibility of MRI morphometric phenotypes in psychiatric disorders.Journal: Nature neuroscienceIn common: Statistics and Machine Learning Toolbox, structural MRI / diffusion, 6 references
- [9] doi:10.1038/s41467-026-73072-6 [code]
- Mapping the spatiotemporal continuum of structural connectivity development across the human connectome in youth.Journal: Nature communicationsIn common: structural MRI / diffusion, 7 references
- [10] doi:10.1523/eneuro.0370-25.2026 [code]
- Pretraining for Large-Scale Functional Connectome Fingerprinting Supports Generalization and Transfer Learning in Functional Neuroimaging.Journal: eNeuroIn common: 7 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 4 scripts, and 2 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:9e9c428ec77e27a7…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
