Biological brain aging, cognitive-motor decline and vascular risk: a multivariate imaging analysis of 40,579 individuals.
The 3 matches
- [1] § Materials and methods › Statistics › Partial least square correlation analysis ↔ pyls/base.py, lines 438–527 · score 0.77 · singular vector, standard error, bootstrap resampling, bootstrap ratio, replacement, weight
- [2] § Materials and methods › Statistics › Partial least square correlation analysis ↔ pyls/structures.py, lines 299–326 · score 0.68 · confidence intervals, singular vector, latent variable, permuting, resampling, correlation
- [3] § Results › Imaging markers of biological brain aging are associated with cognitive and motor function ↔ pyls/types/behavioral.py, lines 172–227 · score 0.59 · confidence interval, cross validation, bootstrap ratio, squares, scores, PLS
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 759 lines · 31 KB · GPL-2.0 · 1 match
- # -*- coding: utf-8 -*-
- import gc
- import warnings
- import numpy as np
- from sklearn.utils.validation import check_random_state
- from . import compute, structures, utils
- def gen_permsamp(groups, n_cond, n_perm, seed=None, verbose=True):
- """
- Generates permutation arrays for PLS permutation testing
- Parameters
- ----------
- groups : (G,) list
- List with number of subjects in each of `G` groups
- n_cond : int
- Number of conditions, for each subject. Default: 1
- n_perm : int
- Number of permutations for which to generate resampling arrays
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- verbose : bool, optional
- Whether to print status updates as permutations are generated.
- Default: True
- Returns
- -------
- permsamp : (S, P) `numpy.ndarray`
- Subject permutation arrays, where `S` is the number of subjects and `P`
- is the requested number of permutations (i.e., `P = n_perm`)
- """
- Y = utils.dummy_code(groups, n_cond)
- permsamp = np.zeros(shape=(len(Y), n_perm), dtype=int)
- subj_inds = np.arange(np.sum(groups), dtype=int)
- rs = check_random_state(seed)
- warned = False
- # calculate some variables for permuting conditions within subject
- # do this here to save on calculation time
- indices, grps = np.where(Y)
- grp_conds = np.split(indices, np.where(np.diff(grps))[0] + 1)
- to_permute = [np.vstack(grp_conds[i:i + n_cond]) for i in
- range(0, Y.shape[-1], n_cond)]
- splitinds = np.cumsum(groups)[:-1]
- check_grps = utils.dummy_code(groups).T.astype(bool)
- for i in utils.trange(n_perm, verbose=verbose, desc='Making permutations'):
- count, duplicated = 0, True
- while duplicated and count < 500:
- count, duplicated = count + 1, False
- # generate conditions permuted w/i subject
- inds = np.hstack([utils.permute_cols(i, seed=rs) for i
- in to_permute])
- # generate permutation of subjects across groups
- perm = rs.permutation(subj_inds)
- # confirm subjects *are* mixed across groups
- if len(groups) > 1:
- for grp in check_grps:
- if np.all(np.sort(perm[grp]) == subj_inds[grp]):
- duplicated = True
- # permute conditions w/i subjects across groups and stack
- perminds = np.hstack([f.flatten('F') for f in
- np.split(inds[:, perm].T, splitinds)])
- # make sure permuted indices are not a duplicate sequence
- dupe_seq = perminds[:, None] == permsamp[:, :i]
- if dupe_seq.all(axis=0).any():
- duplicated = True
- # if we broke out because we tried 500 permutations and couldn't
- # generate a new one, just warn that we're using duplicate
- # permutations and give up
- if count == 500 and not warned:
- warnings.warn('WARNING: Duplicate permutations used.')
- warned = True
- # store the permuted indices
- permsamp[:, i] = perminds
- return permsamp
- def gen_bootsamp(groups, n_cond, n_boot, seed=None, verbose=True):
- """
- Generates bootstrap arrays for PLS bootstrap resampling
- Parameters
- ----------
- groups : (G,) list
- List with number of subjects in each of `G` groups
- n_cond : int
- Number of conditions, for each subject. Default: 1
- n_boot : int
- Number of boostraps for which to generate resampling arrays
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- verbose : bool, optional
- Whether to print status updates as bootstrap samples are genereated.
- Default: True
- Returns
- -------
- bootsamp : (S, B) `numpy.ndarray`
- Subject bootstrap arrays, where `S` is the number of subjects and `B`
- is the requested number of bootstraps (i.e., `B = n_boot`)
- """
- Y = utils.dummy_code(groups, n_cond)
- bootsamp = np.zeros(shape=(len(Y), n_boot), dtype=int)
- subj_inds = np.arange(np.sum(groups), dtype=int)
- rs = check_random_state(seed)
- warned = False
- min_subj = int(np.ceil(Y.sum(axis=0).min() * 0.5))
- # calculate some variables for ensuring we resample with replacement
- # subjects across all their conditions. do this here to save on
- # calculation time
- indices, grps = np.where(Y)
- grp_conds = np.split(indices, np.where(np.diff(grps))[0] + 1)
- inds = np.hstack([np.vstack(grp_conds[i:i + n_cond]) for i
- in range(0, len(grp_conds), n_cond)])
- splitinds = np.cumsum(groups)[:-1]
- check_grps = utils.dummy_code(groups).T.astype(bool)
- for i in utils.trange(n_boot, verbose=verbose, desc='Making bootstraps'):
- count, duplicated = 0, True
- while duplicated and count < 500:
- count, duplicated = count + 1, False
- # empty container to store current bootstrap attempt
- boot = np.zeros(shape=(subj_inds.size), dtype=int)
- # iterate through and resample from w/i groups
- for grp in check_grps:
- curr_grp, all_same = subj_inds[grp], True
- while all_same:
- num_subj = curr_grp.size
- boot[curr_grp] = np.sort(rs.choice(curr_grp,
- size=num_subj,
- replace=True),
- axis=0)
- # make sure bootstrap has enough unique subjs
- if np.unique(boot[curr_grp]).size >= min_subj:
- all_same = False
- # resample subjects (with conditions) and stack groups
- bootinds = np.hstack([f.flatten('F') for f in
- np.split(inds[:, boot].T, splitinds)])
- # make sure bootstrap is not a duplicated sequence
- for grp in check_grps:
- curr_grp = subj_inds[grp]
- check = bootinds[curr_grp, None] == bootsamp[curr_grp, :i]
- if check.all(axis=0).any():
- duplicated = True
- # if we broke out because we tried 500 bootstraps and couldn't
- # generate a new one, just warn that we're using duplicate
- # bootstraps and give up
- if count == 500 and not warned:
- warnings.warn('WARNING: Duplicate bootstraps used.')
- warned = True
- # store the bootstrapped indices
- bootsamp[:, i] = bootinds
- return bootsamp
- def gen_splits(groups, n_cond, n_split, seed=None, test_size=0.5):
- """
- Generates splitting arrays for PLS split-half resampling and CV
- Parameters
- ----------
- groups : (G,) list
- List with number of subjects in each of `G` groups
- n_cond : int
- Number of conditions, for each subject. Default: 1
- n_split : int
- Number of splits for which to generate resampling arrays
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- test_size : (0, 1) float, optional
- Percent of subjects to include in the split halves. Default: 0.5
- Returns
- -------
- splitsamp : (S, I) `numpy.ndarray`
- Subject split arrays, where `S` is the number of subjects and `I`
- is the requested number of splits (i.e., `I = n_split`)
- """
- Y = utils.dummy_code(groups, n_cond)
- splitsamp = np.zeros(shape=(len(Y), n_split), dtype=bool)
- subj_inds = np.arange(np.sum(groups), dtype=int)
- rs = check_random_state(seed)
- warned = False
- # calculate some variables for permuting conditions within subject
- # do this here to save on calculation time
- indices, grps = np.where(Y)
- grp_conds = np.split(indices, np.where(np.diff(grps))[0] + 1)
- inds = np.hstack([np.vstack(grp_conds[i:i + n_cond]) for i
- in range(0, len(grp_conds), n_cond)])
- splitinds = np.cumsum(groups)[:-1]
- check_grps = utils.dummy_code(groups).T.astype(bool)
- for i in range(n_split):
- count, duplicated = 0, True
- while duplicated and count < 500:
- count, duplicated = count + 1, False
- # empty containter to store current split half attempt
- split = np.zeros(shape=(subj_inds.size), dtype=bool)
- # iterate through and split each group separately
- for grp in check_grps:
- curr_grp = subj_inds[grp]
- take = rs.choice([np.ceil, np.floor])
- num_subj = int(take(curr_grp.size * (1 - test_size)))
- splinds = rs.choice(curr_grp,
- size=num_subj,
- replace=False)
- split[splinds] = True
- # split subjects (with conditions) and stack groups
- half = np.hstack([f.flatten('F') for f in
- np.split(((inds + 1).astype(bool)
- * [split[None]]).T,
- splitinds)])
- # make sure split half is not a duplicated sequence
- dupe_seq = half[:, None] == splitsamp[:, :i]
- if dupe_seq.all(axis=0).any():
- duplicated = True
- if count == 500 and not warned:
- warnings.warn('WARNING: Duplicate split halves used.')
- warned = True
- splitsamp[:, i] = half
- return splitsamp
- class BasePLS():
- """
- Base PLS class to be subclassed
- Contains most of the math required for PLS, leaving a few functions for PLS
- subclasses to implement. This will not run without those implementations.
- Parameters
- ----------
- {input_matrix}
- {groups}
- {conditions}
- **kwargs : optional
- Additional key-value pairs; see :obj:`pyls.structures.PLSInputs` for
- more info
- References
- ----------
- {references}
- """.format(**structures._pls_input_docs)
- def __init__(self, X, Y=None, groups=None, n_cond=1, **kwargs):
- # if groups aren't provided or are provided wrong, fix them
- if groups is None:
- groups = [len(X) // n_cond]
- elif not isinstance(groups, (list, np.ndarray)):
- groups = [groups]
- # coerce groups to integers
- groups = [int(g) for g in groups]
- # check that data matrices and groups + n_cond inputs jibe
- n_samples = sum([g * n_cond for g in groups])
- if len(X) != n_samples:
- raise ValueError('Number of samples specified by `groups` and '
- '`n_cond` does not match number of samples in '
- 'input array(s).\n'
- ' EXPECTED: {}\n'
- ' ACTUAL: {} (groups: {} * n_cond: {})'
- .format(len(X), n_samples, groups, n_cond))
- if Y is not None and len(X) != len(Y):
- raise ValueError('Provided `X` and `Y` matrices must have the '
- 'same number of samples. Provided matrices '
- 'differed: X: {}, Y: {}'.format(len(X), len(Y)))
- self.inputs = structures.PLSInputs(X=X, Y=Y, groups=groups,
- n_cond=n_cond, **kwargs)
- # store dummy-coded array of groups / conditions (save on computation)
- self.dummy = utils.dummy_code(groups, n_cond)
- self.rs = check_random_state(self.inputs.get('seed'))
- # check for parallel processing desire
- n_proc = self.inputs.get('n_proc')
- if n_proc is not None and n_proc != 1 and not utils.joblib_avail:
- self.inputs.n_proc = None
- warnings.warn('Setting n_proc > 1 requires the joblib module. '
- 'Considering installing joblib and re-running this '
- 'if you would like parallelization. Resetting '
- 'n_proc to 1 for now.')
- def gen_covcorr(self, X, Y, groups=None):
- """
- Should generate cross-covariance array to be used in `self._svd()`
- Must accept the listed parameters and return one array
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- groups : (G,) array_like
- Array with number of subjects in each of `G` groups
- Returns
- -------
- crosscov : np.ndarray
- Covariance array for decomposition
- """
- raise NotImplementedError
- def gen_distrib(self, X, Y, groups=None, original=None):
- """
- Should generate behavioral correlations or contrast for bootstrap
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- groups : (S, J) array_like
- Dummy coded array, where `S` is observations and `J` corresponds to
- the number of different groups x conditions represented in `X` and
- `Y`. A value of 1 indicates that an observation belongs to a
- specific group or condition
- Returns
- -------
- distrib : (T, L)
- Behavioral correlations or contrast for single bootstrap resample
- """
- raise NotImplementedError
- def run_pls(self, X, Y):
- """
- Runs PLS analysis
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- Returns
- -------
- results : :obj:`pyls.structures.PLSResults`
- Results of PLS (not including PLS type-specific outputs)
- """
- # initate results structure
- self.res = res = structures.PLSResults(inputs=self.inputs)
- # get original singular vectors / values
- res['x_weights'], res['singvals'], res['y_weights'] = \
- self.svd(X, Y, seed=self.rs)
- res['x_scores'] = X @ res['x_weights']
- if self.inputs.n_perm > 0:
- # compute permutations and get statistical significance of LVs
- d_perm, ucorrs, vcorrs = self.permutation(X, Y, seed=self.rs)
- res['permres']['pvals'] = compute.perm_sig(res['singvals'], d_perm)
- res['permres']['permsamples'] = self.permsamp
- if self.inputs.n_split is not None:
- # get ucorr / vcorr (via split half resampling) for original,
- # unpermuted `X` and `Y` arrays
- di = np.linalg.inv(res['singvals'])
- orig_ucorr, orig_vcorr = self.split_half(X, Y,
- res['x_weights'] @ di,
- res['y_weights'] @ di,
- seed=self.rs)
- # get p-values for ucorr/vcorr
- ucorr_prob = compute.perm_sig(np.diag(orig_ucorr), ucorrs)
- vcorr_prob = compute.perm_sig(np.diag(orig_vcorr), vcorrs)
- # get confidence intervals for ucorr/vcorr
- ucorr_ll, ucorr_ul = compute.boot_ci(ucorrs, ci=self.inputs.ci)
- vcorr_ll, vcorr_ul = compute.boot_ci(vcorrs, ci=self.inputs.ci)
- # update results object with split-half resampling results
- res['splitres'].update(dict(ucorr=orig_ucorr,
- vcorr=orig_vcorr,
- ucorr_pvals=ucorr_prob,
- vcorr_pvals=vcorr_prob,
- ucorr_lolim=ucorr_ll,
- vcorr_lolim=vcorr_ll,
- ucorr_uplim=ucorr_ul,
- vcorr_uplim=vcorr_ul))
- return res
- def svd(self, X, Y, groups=None, seed=None):
- """
- Runs SVD on cross-covariance matrix computed from `X` and `Y`
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- groups : (S, J) array_like
- Dummy coded array, where `S` is observations and `J` corresponds to
- the number of different groups x conditions represented in `X` and
- `Y`. A value of 1 indicates that an observation belongs to a
- specific group or condition
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- Returns
- -------
- U : (B, L) `numpy.ndarray`
- Left singular vectors from singular value decomposition
- d : (L, L) `numpy.ndarray`
- Diagonal array of singular values from singular value decomposition
- V : (J, L) `numpy.ndarray`
- Right singular vectors from singular value decomposition
- """
- # make dummy-coded grouping array if not provided
- if groups is None:
- groups = utils.dummy_code(self.inputs.groups, self.inputs.n_cond)
- # generate cross-covariance matrix and determine # of components
- crosscov = self.gen_covcorr(X, Y, groups=groups)
- U, d, V = compute.svd(crosscov, seed=seed)
- return U, d, V
- def bootstrap(self, X, Y, seed=None):
- """
- Bootstraps `X` and `Y` (w/replacement) and recomputes SVD
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- Returns
- -------
- distrib : (T, L) numpy.ndarray
- Either behavioral correlations or group x condition contrast;
- depends on PLS type
- u_sum : (B, L) numpy.ndarray
- Sum of the left singular vectors across all bootstraps
- u_square : (B, L) numpy.ndarray
- Sum of the squared left singular vectors across all bootstraps
- """
- # generate bootstrap resampled indices (unless already provided)
- self.bootsamp = self.inputs.get('bootsamples', None)
- if self.bootsamp is None:
- self.bootsamp = gen_bootsamp(self.inputs.groups,
- self.inputs.n_cond,
- self.inputs.n_boot,
- seed=seed,
- verbose=self.inputs.verbose)
- # make empty arrays to store bootstrapped singular vectors
- # these will be used to calculate the standard error later on for
- # creation of bootstrap ratios
- u_sum = np.zeros_like(self.res['x_weights'])
- u_square = np.zeros_like(self.res['x_weights'])
- # `distrib` corresponds either to the behavioral correlations (if
- # running a behavioral PLS) or to the group/condition contrast (if
- # running a mean-centered PLS); we'll just extend it and then stack
- # all the individual matrices together later (they're quite small so we
- # don't need to be too worried about memory usage, here)
- distrib = []
- # determine the number of bootstraps we'll run each iteration
- iters = 1 if self.inputs.n_proc is None else self.inputs.n_proc
- gen = utils.trange(self.inputs.n_boot, verbose=self.inputs.verbose,
- desc='Running bootstraps')
- with utils.get_par_func(self.inputs.n_proc,
- self.__class__._single_boot) as (par, func):
- boots = 0
- while boots < self.inputs.n_boot:
- # determine number of bootstraps to run this round
- # we don't want to overshoot the requested number, so make
- # sure to cut it off if that's what wold happen
- top = boots + iters
- if top >= self.inputs.n_boot:
- top = self.inputs.n_boot
- # run the bootstraps
- d, usu = zip(*par(func(self, X=X, Y=Y,
- inds=self.bootsamp[..., i],
- groups=self.dummy,
- original=self.res['x_weights'],
- seed=i)
- for i in range(boots, top)))
- # sum bootstrapped singular vectors and store
- u_sum += np.sum(usu, axis=0)
- u_square += np.sum(np.square(usu), axis=0)
- distrib.extend(d)
- # force garbage collection
- # this is only really needed when parallelizing bootstraps
- # the `usu` variable can get REALLY GIANT if either `X` or `Y`
- # is large and `n_proc` is > 1, so we really don't want to keep
- # it around for any longer than absolutely necessary
- if self.inputs.n_proc is not None:
- del usu
- gc.collect()
- # update progress bar and # of bootstraps already run
- gen.update(top - boots)
- boots = top
- gen.close()
- return np.stack(distrib, axis=-1), u_sum, u_square
- def _single_boot(self, X, Y, inds, groups=None, original=None, seed=None):
- """
- Bootstraps `X` and `Y` (w/replacement) and recomputes SVD
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- groups : (S, J) array_like
- Dummy coded input array, where `S` is observations and `J`
- corresponds to the number of different groups x conditions. A value
- of 1 indicates that an observation belongs to a specific group or
- condition.
- original : (B, L) array_like
- Left singular vector from original decomposition of `X` and `Y`.
- Used to perform Procrustes rotation on permuted singular vectors
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- Returns
- -------
- distrib : np.ndarray
- Either behavioral correlations or contrast, depending on PLS type;
- generated with self.gen_distrib() which should be specified by the
- PLS subclass
- U_sum : (B, L) array_like
- Left singular vectors from decomposition of bootstrap resampled `X`
- and `Y`
- """
- # make sure we have original (non-bootstrapped) singular vectors
- # these are required for the procrustes rotation to ensure our
- # singular vectors are all in the same orientation
- if original is None:
- original = self.svd(X, Y, groups=groups, seed=seed)[0]
- # perform SVD of bootstrapped arrays and rotate left singular vectors
- U, d = self.svd(X[inds], Y[inds], groups=groups, seed=seed)[:-1]
- U_boot = compute.procrustes(original, U, d)
- # get contrast / behavcorrs (this function should be specified by the
- # subclass)
- distrib = self.gen_distrib(X[inds], Y[inds], original, groups)
- return distrib, U_boot
- def make_permutation(self, X, Y, perminds):
- """
- Permutes `Y` according to `perminds`, leaving `X` un-permuted
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- perminds : (S,) array_like
- Array by which to permute `Y`
- Returns
- -------
- Xp : (S, B) array_like
- Identical to `X`
- Yp : (S, T) array_like
- `Y`, permuted according to `perminds`
- """
- return X, Y[perminds]
- def permutation(self, X, Y, seed=None):
- """
- Permutes `X` (w/o replacement) and recomputes SVD
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- Returns
- -------
- d_perm : (L, P) `numpy.ndarray`
- Permuted singular values, where `L` is the number of singular
- values and `P` is the number of permutations
- ucorrs : (L, P) `numpy.ndarray`
- Split-half correlations of left singular values. Only set if
- `self.inputs.n_split != 0`
- vcorrs : (L, P) `numpy.ndarray`
- Split-half correlations of right singular values. Only set if
- `self.inputs.n_split != 0`
- """
- # generate permuted indices (unless already provided)
- self.permsamp = self.inputs.get('permsamples')
- if self.permsamp is None:
- self.permsamp = gen_permsamp(self.inputs.groups,
- self.inputs.n_cond,
- self.inputs.n_perm,
- seed=seed,
- verbose=self.inputs.verbose)
- # get permuted values (parallelizing as requested)
- gen = utils.trange(self.inputs.n_perm, verbose=self.inputs.verbose,
- desc='Running permutations')
- with utils.get_par_func(self.inputs.n_proc,
- self.__class__._single_perm) as (par, func):
- out = par(func(self, X=X, Y=Y, inds=self.permsamp[:, i],
- groups=self.dummy, original=self.res['y_weights'],
- seed=i)
- for i in gen)
- d_perm, ucorrs, vcorrs = [np.stack(o, axis=-1) for o in zip(*out)]
- return d_perm, ucorrs, vcorrs
- def _single_perm(self, X, Y, inds, groups=None, original=None, seed=None):
- """
- Permutes `X` (w/o replacement) and recomputes SVD
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- inds : (S,) array_like
- Permutation resampling array
- original : (J, L) array_like
- Right singular vector from original decomposition of `X` and `Y`.
- Used to perform Procrustes rotation on permuted singular values,
- if desired
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- Returns
- -------
- d_perm : (L,) `numpy.ndarray`
- Permuted singular values, where `L` is the number of singular
- values
- ucorrs : (L,) `numpy.ndarray`
- Split-half correlations of left singular values. Only set if
- `self.inputs.n_split != 0`
- vcorrs : (L,) `numpy.ndarray`
- Split-half correlations of right singular values. Only set if
- `self.inputs.n_split != 0`
- """
- # calculate SVD of permuted matrices
- Xp, Yp = self.make_permutation(X, Y, inds)
- U, d, V = self.svd(Xp, Yp, groups=groups, seed=seed)
- # optionally get rotated/rescaled singular values
- if self.inputs.rotate:
- if original is None:
- original = self.svd(X, Y, groups=groups, seed=seed)[-1]
- ssd = np.sqrt(np.sum(compute.procrustes(original, V, d)**2,
- axis=0))
- else:
- ssd = np.diag(d)
- # get ucorr/vcorr if split-half resampling requested
- if self.inputs.n_split is not None:
- di = np.linalg.inv(d)
- ucorr, vcorr = self.split_half(Xp, Yp, U @ di, V @ di,
- groups=groups, seed=seed)
- else:
- ucorr, vcorr = None, None
- return ssd, ucorr, vcorr
- def split_half(self, X, Y, ud=None, vd=None, groups=None, seed=None):
- """
- Parameters
- ----------
- X : (S, B) array_like
- Input data matrix, where `S` is observations and `B` is features
- Y : (S, T) array_like
- Input data matrix, where `S` is observations and `T` is features
- ud : (B, L) array_like
- Left singular vectors, scaled by singular values
- vd : (J, L) array_like
- Right singular vectors, scaled by singular values
- seed : {int, :obj:`numpy.random.RandomState`, None}, optional
- Seed for random number generation. Default: None
- Returns
- -------
- ucorr : (L,) `numpy.ndarray`
- Average correlation of left singular vectors across split-halves
- vcorr : (L,) `numpy.ndarray`
- Average correlation of right singular vectors across split-halves
- """
- # generate splits
- splitsamp = gen_splits(self.inputs.groups,
- self.inputs.n_cond,
- self.inputs.n_split,
- seed=seed,
- test_size=0.5).astype(bool)
- # make dummy-coded grouping array if not provided
- if groups is None:
- groups = utils.dummy_code(self.inputs.groups, self.inputs.n_cond)
- # generate original singular vectors if not provided
- if ud is None or vd is None:
- U, d, V = self.svd(X, Y, groups=groups, seed=seed)
- di = np.linalg.inv(d)
- ud, vd = U @ di, V @ di
- # empty arrays to hold split-half correlations
- ucorr = np.zeros(shape=(ud.shape[-1], self.inputs.n_split))
- vcorr = np.zeros(shape=(vd.shape[-1], self.inputs.n_split))
- for i in range(self.inputs.n_split):
- # calculate cross-covariance matrix for both splits
- spl = splitsamp[:, i]
- D1 = self.gen_covcorr(X[spl], Y[spl], groups=groups[spl])
- D2 = self.gen_covcorr(X[~spl], Y[~spl], groups=groups[~spl])
- # project cross-covariance matrices onto original SVD to obtain
- # left & right singular vector and correlate between split halves
- ucorr[:, i] = compute.efficient_corr(D1.T @ vd, D2.T @ vd)
- vcorr[:, i] = compute.efficient_corr(D1 @ ud, D2 @ ud)
- # return average correlations for singular vectors across `n_split`
- return np.mean(ucorr, axis=-1), np.mean(vcorr, axis=-1)
base.py at commit d8a19d5, under GPL-2.0 · at the source
Overview
- Department of Neurology, University Medical Center Hamburg-Eppendorf, Hamburg, Germany
- Department of Diagnostic and Interventional Neuroradiology, University Medical Center Hamburg-Eppendorf, Hamburg, Germany
- Department of Psychiatry and Psychotherapy, University Medical Center Hamburg-Eppendorf, Hamburg, Germany
- Department of General and Interventional Cardiology, University Heart and Vascular Center, Hamburg, Germany
- Epidemiological Study Center, University Medical Center Hamburg-Eppendorf, Hamburg, Germany
- German Center for Cardiovascular Research (DZHK), Partner Site Hamburg/Kiel/Luebeck, Hamburg, Germany
- University Center of Cardiovascular Science, University Heart and Vascular Center, Hamburg, Germany
- Faculty of Medicine, Institute of Systems Neuroscience, Heinrich Heine University Düsseldorf, Düsseldorf, Germany
- Institute of Neuroscience and Medicine, Brain and Behaviour (INM-7), Research Center Jülich, Jülich, Germany
Abstract
Introduction: Age-related declines in cognitive and motor functions show highly variable trajectories. To better understand the underlying mechanisms, we investigated multivariate associative effects between modifiable vascular risk factors, biological brain aging, cognitive, and motor performance in 40,579 individuals from the population-based UK Biobank and Hamburg City Health Study.
Methods: We employed partial least squares correlation analysis (PLS) to model associations between multi-domain cognitive and motor test scores and three distinct MRI-derived markers of biological brain aging: relative brain age (from morphometric brain imaging), white matter hyperintensity load, and peak width of skeletonized mean diffusivity. Furthermore, we conducted mediation analyses to assess if these markers mediate the impact of vascular risk on functional decline.
Results: PLS identified a single dominant latent dimension explaining 94.7% of the shared variance between neuroimaging and behavior. This dimension linked higher biological brain aging markers – with relative brain age showing the strongest contribution – to poorer cognitive and motor performance, particularly in executive function and processing speed. Mediation analysis revealed that biological brain aging acts as a partial mediator for the negative effects of blood pressure, glucose, waist-hip ratio, and smoking load on cognitive and motor function. Notably, this mediating effect was not observed for cholesterol levels. These results were consistent across both cohorts.
Discussion: Our study illustrates the associative interplay between vascular health, biological brain aging, and cognitive and motor performance, emphasizing the need for preventive strategies to maintain late-life independence in aging populations.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.
rmarkello/pyls
d8a19d564cc5804249527b68c937df3a5fd8c7cc, 4 November 2019Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
34 files
- docs/
conf.py , Python, 110 lines - pyls/
__init__.py , Python, 19 lines - pyls/
_version.py , Python, 520 lines - pyls/
base.py , Python, 759 lines, 1 match - pyls/
compute.py , Python, 414 lines - pyls/
examples/ , Python, 3 lines__init__.py - pyls/
examples/ , Python, 184 linesdatasets.py - pyls/
io.py , Python, 122 lines - pyls/
matlab/ , Python, 8 lines__init__.py - pyls/
matlab/ , Python, 225 linesio.py - pyls/
plotting/ , Python, 183 linesmeancentered.py - pyls/
structures.py , Python, 345 lines, 1 match - pyls/
tests/ , Python, 3 lines__init__.py - pyls/
tests/ , Python, 40 linesconftest.py - pyls/
tests/ , Python, 285 linesmatlab.py - pyls/
tests/ , Python, 180 linestest_base.py - pyls/
tests/ , Python, 49 linestest_compute.py - pyls/
tests/ , Python, 90 linestest_examples.py - pyls/
tests/ , Python, 17 linestest_io.py - pyls/
tests/ , Python, 35 linestest_matlab.py - pyls/
tests/ , Python, 58 linestest_structures.py - pyls/
tests/ , Python, 156 linestest_utils.py - pyls/
tests/ , Python, 1 linetypes/ __init__.py - pyls/
tests/ , Python, 104 linestypes/ test_regression.py - pyls/
tests/ , Python, 143 linestypes/ test_svd.py - pyls/
types/ , Python, 10 lines__init__.py - pyls/
types/ , Python, 294 lines, 1 matchbehavioral.py - pyls/
types/ , Python, 248 linesmeancentered.py - pyls/
types/ , Python, 490 linesregression.py - pyls/
utils.py , Python, 279 lines - setup.py, Python, 14 lines
- versioneer.py, Python, 1,822 lines
- LICENSE, License, 339 lines
- README.md, Text, 124 lines
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 32 scripts, each with its path and the digest of its content;
- 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- ukbiobank.ac.uk/
media/ , at UK Biobank; found in the end of the paper0xsbmfmw
Data availability statement
UK Biobank data can be obtained via its standardized data access procedure (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 17 authors, 7 keywords, 54 references.
Cite
This paper
Petersen, M., Link, M. A., Mayer, C., Nägele, F. L., Schell, M., Jensen, M., Schlemm, E., Fiehler, J., Gallinat, J., Kühn, S., Twerenbold, R., Omidvarnia, A., Hoffstaedter, F., Patil, K. R., Eickhoff, S. B., Thomalla, G., & Cheng, B. (2026). Biological brain aging, cognitive-motor decline and vascular risk: a multivariate imaging analysis of 40,579 individuals. Frontiers in aging neuroscience, 18, 1789408. https://
BibTeX
@article{petersen2026bio
author = {Petersen, Marvin and Link, Moritz A. and Mayer, Carola and Nägele, Felix L. and Schell, Maximilian and Jensen, Märit and Schlemm, Eckhard and Fiehler, Jens and Gallinat, Jürgen and Kühn, Simone and Twerenbold, Raphael and Omidvarnia, Amir and Hoffstaedter, Felix and Patil, Kaustubh R. and Eickhoff, Simon B. and Thomalla, Götz and Cheng, Bastian},
title = {{Biological brain aging, cognitive-motor decline and vascular risk: a multivariate imaging analysis of 40,579 individuals}},
journal = {Frontiers in aging neuroscience},
year = {2026},
month = may,
volume = {18},
pages = {1789408},
publisher = {Frontiers Media SA},
issn = {1663-4365},
doi = {10.3389/
url = {https://
pmid = {42147461},
pmcid = {PMC13176138}
}
RIS
TY - JOUR
AU - Petersen, Marvin
AU - Link, Moritz A.
AU - Mayer, Carola
AU - Nägele, Felix L.
AU - Schell, Maximilian
AU - Jensen, Märit
AU - Schlemm, Eckhard
AU - Fiehler, Jens
AU - Gallinat, Jürgen
AU - Kühn, Simone
AU - Twerenbold, Raphael
AU - Omidvarnia, Amir
AU - Hoffstaedter, Felix
AU - Patil, Kaustubh R.
AU - Eickhoff, Simon B.
AU - Thomalla, Götz
AU - Cheng, Bastian
TI - Biological brain aging, cognitive-motor decline and vascular risk: a multivariate imaging analysis of 40,579 individuals
T2 - Frontiers in aging neuroscience
J2 - Front Aging Neurosci
PY - 2026
DA - 2026/
VL - 18
SP - 1789408
SN - 1663-4365
PB - Frontiers Media SA
DO - 10.3389/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3389/
"type": "article-journal",
"title": "Biological brain aging, cognitive-motor decline and vascular risk: a multivariate imaging analysis of 40,579 individuals",
"container-title": "Frontiers in aging neuroscience",
"author": [
{
"family": "Petersen",
"given": "Marvin"
},
{
"family": "Link",
"given": "Moritz A."
},
{
"family": "Mayer",
"given": "Carola"
},
{
"family": "Nägele",
"given": "Felix L."
},
{
"family": "Schell",
"given": "Maximilian"
},
{
"family": "Jensen",
"given": "Märit"
},
{
"family": "Schlemm",
"given": "Eckhard"
},
{
"family": "Fiehler",
"given": "Jens"
},
{
"family": "Gallinat",
"given": "Jürgen"
},
{
"family": "Kühn",
"given": "Simone"
},
{
"family": "Twerenbold",
"given": "Raphael"
},
{
"family": "Omidvarnia",
"given": "Amir"
},
{
"family": "Hoffstaedter",
"given": "Felix"
},
{
"family": "Patil",
"given": "Kaustubh R."
},
{
"family": "Eickhoff",
"given": "Simon B."
},
{
"family": "Thomalla",
"given": "Götz"
},
{
"family": "Cheng",
"given": "Bastian"
}
],
"container-title-short":
"volume": "18",
"page": "1789408",
"DOI": "10.3389/
"PMID": "42147461",
"PMCID": "PMC13176138",
"ISSN": "1663-4365",
"publisher": "Frontiers Media SA",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1371/journal.pbio.3003856 [code]
- Aging and metabolism contribute separately to brain-body health.Journal: PLoS biologyIn common: seaborn, scikit-learn, pandas, 2 other tools, ukbiobank.ac.uk/media/0xsbmfmw, structural MRI / diffusion, 6 references
- [2] doi:10.1038/s41467-026-71271-9 [code]
- Exposome-wide patterns predict brain health in aging.Journal: Nature communicationsIn common: seaborn, scikit-learn, pandas, 2 other tools, 5 references
- [3] doi:10.7554/elife.108109 [code]
- Multimodal MRI marker of cognition explains the association between cognition and mental health in the UK Biobank.Journal: eLifeIn common: seaborn, scikit-learn, pandas, 2 other tools, cognitive, structural MRI / diffusion, 5 references
- [4] doi:10.1016/j.ebiom.2026.106312 [code]
- Subgingival microbiota composition is associated with brain health in the general population-the PAROMIND study.Journal: EBioMedicineIn common: scikit-learn, pandas, SciPy, 1 other tool, cognitive, 4 references
- [5] doi:10.3389/frai.2026.1771088 [code]
- Few-shot deployment of pretrained MRI transformers in brain imaging tasks.Journal: Frontiers in artificial intelligenceIn common: h5py, seaborn, scikit-learn, 3 other tools, structural MRI / diffusion, 4 references
- [6] doi:10.1038/s41514-026-00456-9 [code]
- Exploring the link between body physiology and cognition: the role of the brain and aging.Journal: npj agingIn common: seaborn, scikit-learn, pandas, 2 other tools, cognitive, 4 references
- [7] doi:10.1186/s40708-026-00316-y [code]
- Generalizable and explainable deep learning for brain MRI: a multi-cohort evaluation of 3D architectures for age and sex prediction.Journal: Brain informaticsIn common: seaborn, scikit-learn, pandas, 2 other tools, structural MRI / diffusion, 3 references
- [8] doi:10.1162/imag.a.1242 [code]
- Stable individual differences dominate adult brain volume variation until later life.Journal: Imaging neuroscience (Cambridge, Mass.)In common: seaborn, pandas, SciPy, 1 other tool, structural MRI / diffusion, 4 references
- [9] doi:10.1038/s41593-026-02359-0 [code]
- The cross-site reproducibility of MRI morphometric phenotypes in psychiatric disorders.Journal: Nature neuroscienceIn common: scikit-learn, pandas, SciPy, 1 other tool, structural MRI / diffusion, 5 references
- [10] doi:10.1016/j.isci.2026.117180 [code]
- Developmental changes in similarity between neural representations of mental arithmetic and artificial neural networks.Journal: iScienceIn common: h5py, seaborn, scikit-learn, 3 other tools, 3 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 32 scripts, and 3 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e015d9f2827857f0…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
