OSCR

Habenular structural-functional dysconnectivity in bipolar disorder: evidence from multimodal imaging and transcriptomic integration.

Code ↔ Paper

4 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 4 matches
  1. [1] § Methods › Associations between functional connectivity and gene expression ↔ abagen/probes_.py, lines 350–379 · score 0.62 · singular vectors, covariance matrix, gene
  2. [2] § Methods › Spatial transcriptome in the brain ↔ abagen/cli/run.py, lines 59–123 · score 0.56 · Allen Human Brain, cortical, abagen, space, probes, donors
  3. [3] § Methods › Spatial transcriptome in the brain ↔ abagen/datasets/fetchers.py, lines 72–160 · score 0.53 · Allen Human Brain, adult, abagen, probes, donors, transcriptome
  4. [4] § Methods › Enrichment of Gene Ontology terms among genes associated with altered functional connectivity involving the habenula ↔ abagen/probes_.py, lines 22–68 · score 0.53 · ENTREZ IDs, Gene symbols

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 763 lines · 31 KB · BSD-3-Clause · 2 matches

  1. # -*- coding: utf-8 -*-
  2. """
  3. Functions for annotating / selecting microarray probes
  4. """
  5. import functools
  6. import gzip
  7. from io import StringIO
  8. import itertools
  9. import logging
  10. from pkg_resources import resource_filename
  11. import numpy as np
  12. import pandas as pd
  13. from scipy import stats as sstats
  14. from . import datasets, io, utils
  15. LGR = logging.getLogger('abagen')
  16. def reannotate_probes(probes):
  17. """
  18. Replaces gene symbols in `probes` with reannotated data
  19. Uses annotations from [PR18]_ to replace probe annotations shipped with
  20. AHBA data. Any probes that were unable to be matched to a gene in
  21. reannotation procedure are not retained.
  22. Parameters
  23. ----------
  24. probes : str or pandas.DataFrame
  25. Probe file or loaded probe dataframe from Allen Brain Institute
  26. containing information on microarray probes
  27. Returns
  28. -------
  29. reannotated : pandas.DataFrame
  30. Provided probe information with updated gene symbols and Entrez IDs
  31. References
  32. ----------
  33. .. [PR18] Arnatkevic̆iūtė, A., Fulcher, B. D., & Fornito, A. (2019). A
  34. practical guide to linking brain-wide gene expression and neuroimaging
  35. data. NeuroImage, 189, 353-367.
  36. """
  37. LGR.info('Reannotating probes with information from Arnatkevic̆iūtė '
  38. 'et al., 2019, NeuroImage')
  39. # load in reannotated probes
  40. reannot = resource_filename('abagen', 'data/reannotated.csv.gz')
  41. with gzip.open(reannot, 'r') as src:
  42. reannot = pd.read_csv(StringIO(src.read().decode('utf-8')))
  43. # merge reannotated with original, keeping only reannotated
  44. probes = io.read_probes(probes).reset_index()
  45. merged = pd.merge(reannot[['probe_name', 'gene_symbol', 'entrez_id']],
  46. probes[['probe_name', 'probe_id']],
  47. on='probe_name', how='left')
  48. # reset index as probe_id and sort
  49. reannotated = merged.set_index('probe_id') \
  50. .sort_index() \
  51. .dropna(subset=['entrez_id'])
  52. reannotated.loc[:, 'entrez_id'] = reannotated['entrez_id'].astype('int')
  53. return reannotated
  54. def filter_probes(pacall, annotation, probes, threshold=0.5):
  55. """
  56. Performs intensity based filtering (IBF) of expression probes
  57. Uses binary indicator for expression levels in `pacall` to determine which
  58. probes have expression levels above background noise in `threshold` of
  59. samples across donors.
  60. Parameters
  61. ----------
  62. pacall : dict
  63. Dictionary where keys are donor IDs and values are filepaths to (or
  64. dataframes of) PACall.csv files from Allen Brain Institute
  65. annotation : dict
  66. Dictionary where keys are donor IDs and values are filepaths to (or
  67. dataframes of) SampleAnnot.csv files from Allen Brain Institute
  68. probes : str or pandas.DataFrame
  69. Filepath to Probes.csv or dataframe containing information on
  70. microarray probes that should be considered in filtering (probes not in
  71. this will be ignored)
  72. threshold : (0, 1) float, optional
  73. Threshold for filtering probes. Specifies the proportion of samples for
  74. which a given probe must have expression levels above background noise.
  75. Default: 0.5
  76. Returns
  77. -------
  78. filtered : pandas.DataFrame
  79. Dataframe containing information on probes that should be retained
  80. according to intensity-based filtering
  81. """
  82. threshold = np.clip(threshold, 0.0, 1.0)
  83. LGR.info(f'Filtering probes with intensity-based threshold of {threshold}')
  84. probes = io.read_probes(probes)
  85. signal, n_samp = np.zeros(len(probes), dtype=int), 0
  86. for donor, pa in pacall.items():
  87. annot = io.read_annotation(annotation[donor]).index
  88. data = io.read_pacall(pa).loc[probes.index, annot]
  89. n_samp += data.shape[-1]
  90. # sum binary expression indicator across samples for current subject
  91. signal += np.asarray(data.sum(axis=1))
  92. # calculate proportion of signal to noise for given probe across samples
  93. keep = (signal / n_samp) >= threshold
  94. LGR.info(f'{keep.sum()} probes survive intensity-based filtering')
  95. return probes[keep]
  96. def _groupby_structure_id(microarray, annotation):
  97. """
  98. Averages samples in `microarray` having identical structure IDs
  99. Parameters
  100. ----------
  101. microarray : (P, S) pandas.DataFrame
  102. Dataframe should have `P` rows representing probes and `S` columns
  103. representing distinct samples, with values indicating microarray
  104. expression levels
  105. annotation : (S, A) pandas.DataFrame
  106. Annotation dataframe, obtained by loading a SampleAnnot.csv file from
  107. Allen Brain Institute
  108. Returns
  109. -------
  110. expression : (P, R) pandas.DataFrame
  111. Input `microarray` dataframe but with `S` samples averaged into `R`
  112. regions
  113. """
  114. sid = io.read_annotation(annotation)['structure_id']
  115. return io.read_microarray(microarray).groupby(sid, axis=1).mean()
  116. def _groupby_and_apply(expression, probes, info, applyfunc):
  117. """
  118. Subsets `expression` based on most representative probe
  119. Parameters
  120. ----------
  121. expression : dict of (P, S) pandas.DataFrame
  122. Dictionary where keys are donor IDs and values are dataframes with `P`
  123. rows representing probes and `S` columns representing distinct samples
  124. probes : pandas.DataFrame
  125. Dataframe containing information on probes that should be considered in
  126. representative analysis. Generally, intensity-based-filtering (i.e.,
  127. `filter_probes()`) should have been used to reduce this list to only
  128. those probes with good expression signal
  129. info : pandas.DataFrame
  130. Dataframe containing information on probe expression information. Index
  131. should be unique probe IDs and must have at least 'gene_symbol' column
  132. applyfunc : callable
  133. Function used to select representative probe ID from those indexing
  134. the same gene. Must accept a pandas dataframe as input and return a
  135. string (i.e., the chosen probe ID)
  136. Returns
  137. -------
  138. representative : dict of (S, G) pandas.DataFrame
  139. Dictionary where keys are donor IDs and values are dataframes with `S`
  140. rows representing distinct samples and `G` columns representing unique
  141. genes
  142. """
  143. # group probes by gene and get probe corresponding to relevant feature
  144. retained = info.groupby('gene_symbol').apply(applyfunc).dropna().squeeze()
  145. probes = probes.loc[sorted(retained)].sort_values('gene_symbol')
  146. # subset expression dataframes to retain only desired probes
  147. representative = {
  148. d: e.loc[probes.index].T
  149. for d, e in utils.check_dict(expression).items()
  150. }
  151. return representative
  152. def _diff_stability(expression, probes, annotation, *args, **kwargs):
  153. """
  154. Picks one probe to represent `expression` data for each gene in `probes`
  155. If there are multiple probes with expression data for the same gene, this
  156. function will calculate the similarity of each probes' expression across
  157. donors and select the probe with the most consistent pattern of regional
  158. variation (i.e., "differential stability" or DS). Regions are defined by
  159. the "structure_id" column in `annotation`; similarity is calculated by the
  160. Spearman correlation coefficient
  161. Parameters
  162. ----------
  163. expression : dict of (P, S) pandas.DataFrame
  164. Dictionary where keys are donor IDs and values are dataframes with `P`
  165. rows representing probes and `S` columns representing distinct samples
  166. probes : pandas.DataFrame
  167. Dataframe containing information on probes that should be considered in
  168. representative analysis. Generally, intensity-based-filtering (i.e.,
  169. `filter_probes()`) should have been used to reduce this list to only
  170. those probes with good expression signal
  171. annotation : list of str
  172. List of filepaths to annotation files from Allen Brain Institute (i.e.,
  173. as obtained by calling :func:`abagen.fetch_microarray` and accessing
  174. the `annotation` attribute on the resulting object).
  175. Returns
  176. -------
  177. representative : dict of (S, G) pandas.DataFrame
  178. Dictionary where keys are donor IDs and values are dataframes with `S`
  179. rows representing distinct samples and `G` columns representing unique
  180. genes
  181. """
  182. # confirm inputs are expected dictionaries
  183. expression = utils.check_dict(expression)
  184. annotation = utils.check_dict(annotation)
  185. # collapse (i.e., average) expression across AHBA anatomical regions
  186. region_exp = [
  187. _groupby_structure_id(microarray, annotation[donor])
  188. for donor, microarray in expression.items()
  189. ]
  190. # get correlation of probe expression across samples for all donor pairs
  191. probe_exp = np.zeros((len(probes), sum(range(len(expression)))))
  192. for n, (exp1, exp2) in enumerate(itertools.combinations(region_exp, 2)):
  193. # samples that current donor pair have in common
  194. samples = np.intersect1d(exp1.columns, exp2.columns)
  195. # the ranking process can take a few seconds on each loop
  196. # unfortunately, we have to do it each time because `samples` changes
  197. # based on which anatomical regions the two subjects have in common
  198. exp1 = exp1.loc[:, samples].T.rank()
  199. exp2 = exp2.loc[:, samples].T.rank()
  200. probe_exp[:, n] = utils.efficient_corr(exp1, exp2)
  201. info = pd.DataFrame(dict(gene_symbol=np.asarray(probes.gene_symbol),
  202. diff_stability=probe_exp.mean(axis=1)),
  203. index=probes.index)
  204. applyfunc = functools.partial(_max_idx, column='diff_stability')
  205. return _groupby_and_apply(expression, probes, info, applyfunc)
  206. def _rnaseq(expression, probes, annotation, *args, **kwargs):
  207. """
  208. Picks one probe to represent `expression` data for each gene in `probes`
  209. If there are multiple probes with expression data for the same gene, this
  210. function will calculate the similarity between each probes' microarray
  211. expression and RNAseq expression data of the relevant gene, selecting the
  212. probe with the greatest similarity to the RNAseq data. Regions are defined
  213. by the "structure_id" column in `annotation`; similarity is calculated by
  214. the Spearman correlation coefficient.
  215. Parameters
  216. ----------
  217. expression : dict of (P, S) pandas.DataFrame
  218. Dictionary where keys are donor IDs and values are dataframes with `P`
  219. rows representing probes and `S` columns representing distinct samples
  220. probes : pandas.DataFrame
  221. Dataframe containing information on probes that should be considered in
  222. representative analysis. Generally, intensity-based-filtering (i.e.,
  223. `filter_probes()`) should have been used to reduce this list to only
  224. those probes with good expression signal
  225. annotation : list of str
  226. List of filepaths to annotation files from Allen Brain Institute (i.e.,
  227. as obtained by calling :func:`abagen.fetch_microarray` and accessing
  228. the `annotation` attribute on the resulting object).
  229. Returns
  230. -------
  231. representative : dict of (S, G) pandas.DataFrame
  232. Dictionary where keys are donor IDs and values are dataframes with `S`
  233. rows representing distinct samples and `G` columns representing unique
  234. genes
  235. """
  236. # confirm inputs are expected dictionaries
  237. expression = utils.check_dict(expression)
  238. annotation = utils.check_dict(annotation)
  239. # fetch RNAseq data
  240. rnaseq = datasets.fetch_rnaseq(donors=expression.keys())
  241. probe_exp = np.ones((len(probes), len(rnaseq))) * np.nan
  242. for n, (donor, data) in enumerate(rnaseq.items()):
  243. # collapse (i.e., average) data across AHAB anatomical regions
  244. micro = _groupby_structure_id(expression[donor], annotation[donor])
  245. rna = _groupby_structure_id(io.read_tpm(data['tpm']),
  246. data['annotation'])
  247. # get rid of "constant" RNAseq genes
  248. rna = rna[np.logical_not(np.isclose(rna.std(axis=1, ddof=1), 0))]
  249. # get matching genes + strcutres between microarray + RNAseq
  250. regions = np.intersect1d(micro.columns, rna.columns)
  251. mask = np.isin(np.asarray(probes.gene_symbol),
  252. np.intersect1d(probes.gene_symbol, rna.index))
  253. genes = np.asarray(probes.loc[mask, 'gene_symbol'])
  254. micro, rna = micro.loc[mask, regions].T, rna.loc[genes, regions].T
  255. # correlate expression values across regions for each gene
  256. probe_exp[mask, n] = utils.efficient_corr(micro.rank(), rna.rank())
  257. mask = np.sum(np.isnan(probe_exp), axis=1) < len(rnaseq)
  258. info = pd.DataFrame(dict(gene_symbol=np.asarray(probes[mask].gene_symbol),
  259. rna_corr=np.nanmean(probe_exp[mask], axis=1)),
  260. index=probes.index[mask])
  261. applyfunc = functools.partial(_max_idx, column='rna_corr')
  262. return _groupby_and_apply(expression, probes, info, applyfunc)
  263. def _max_idx(df, column):
  264. """
  265. Returns probe ID with max index in `df`
  266. Parameters
  267. ----------
  268. df : (P, 1) pandas.DataFrame
  269. Dataframe with `P` rows indicating distinct probes and one column
  270. containing summary statistic of probe expression
  271. column : str, optional
  272. Column name from which to extract the max index. If not specified uses
  273. the first numerical column.
  274. Returns
  275. -------
  276. probe_id : str
  277. ID of probe selected as representative for given gene
  278. """
  279. return df.idxmax(numeric_only=True)[column]
  280. def _max_loading(df):
  281. """
  282. Returns probe ID with max loading along first principal component of `df`
  283. Parameters
  284. ----------
  285. df : (P, S) pandas.DataFrame
  286. Dataframe with `P` rows indicating distinct probe expression values
  287. across `S` samples
  288. Returns
  289. -------
  290. probe_id : str
  291. ID of probe selected as representative for given gene
  292. """
  293. if len(df) == 1:
  294. return df.index[0]
  295. data = np.asarray(df)
  296. data = data - data.mean(axis=0, keepdims=True)
  297. # svd() is faster than eig() here because we don't need to construct the
  298. # covariance matrix
  299. u, s, v = np.linalg.svd(data, full_matrices=False)
  300. # use sign flip based on right singular vectors (as we would with eig())
  301. v *= np.sign(v[range(len(v)), np.argmax(np.abs(v), axis=1)])[:, np.newaxis]
  302. return df.index[(data @ v.T)[:, 0].argmax()]
  303. def _correlate(df, method):
  304. """
  305. Returns probe ID with max avg correlation (>2 probes) or `method` (<2)
  306. Parameters
  307. ----------
  308. df : (P, S) pandas.DataFrame
  309. Dataframe with `P` rows indicating distinct probe expression values
  310. across `S` samples
  311. method : {'variance', 'intensity'}
  312. Method for selecting representative probe when only two probes index
  313. the same gene (>2 probes uses correlation method)
  314. Returns
  315. -------
  316. probe_id : str
  317. ID of probe selected as representative for given gene
  318. """
  319. data = np.asarray(df)
  320. if len(data) > 2:
  321. xmax = np.mean((np.corrcoef(data) + 1) / 2, axis=1)
  322. elif len(data) == 2:
  323. if method == 'variance':
  324. xmax = np.std(data, axis=1)
  325. elif method == 'intensity':
  326. xmax = np.mean(data, axis=1)
  327. else:
  328. return df.index[0]
  329. return df.index[xmax.argmax()]
  330. def _average(expression, probes, *args, **kwargs):
  331. """
  332. Averages expression data for probes representing the same gene
  333. Parameters
  334. ----------
  335. expression : dict of (P, S) pandas.DataFrame
  336. Dictionary where keys are donor IDs and values are dataframes with `P`
  337. rows representing probes and `S` columns representing distinct samples
  338. probes : pandas.DataFrame
  339. Dataframe containing information on microarray probes that should be
  340. considered in representative analysis. Generally intensity-based
  341. filtering (i.e., via :func:`filter_probes()`) should have been used to
  342. reduce this list to only those probes with good expression signal
  343. Returns
  344. -------
  345. representative : dict of (S, G) pandas.DataFrame
  346. Dictionary where keys are donor IDs and values are dataframes with `S`
  347. rows representing distinct samples and `G` columns representing unique
  348. genes
  349. """
  350. def _avg(df):
  351. return df.rename(probes['gene_symbol'].to_dict()) \
  352. .rename_axis('gene_symbol') \
  353. .groupby('gene_symbol') \
  354. .mean().T
  355. return {d: _avg(exp) for d, exp in utils.check_dict(expression).items()}
  356. def _collapse(expression, probes, *args, method='max_variance', **kwargs):
  357. """
  358. Selects one representative probe per gene using provided `method`
  359. Parameters
  360. ----------
  361. expression : dict of (P, S) pandas.DataFrame
  362. Dictionary where keys are donor IDs and values are dataframes with `P`
  363. rows representing probes and `S` columns representing distinct samples
  364. probes : pandas.DataFrame
  365. Dataframe containing information on microarray probes that should be
  366. considered in representative analysis. Generally intensity-based
  367. filtering (i.e., via :func:`filter_probes()`) should have been used to
  368. reduce this list to only those probes with good expression signal
  369. method : str, optional
  370. Method by which to select represenative probes for each gene. Must be
  371. one of ['max_variance', 'max_intensity', 'pc_loading', 'corr_variance',
  372. 'corr_intensity']
  373. Returns
  374. -------
  375. representative : dict of (S, G) pandas.DataFrame
  376. Dictionary where keys are donor IDs and values are dataframes with `S`
  377. rows representing distinct samples and `G` columns representing unique
  378. genes
  379. """
  380. # concatenate all donors into giant probe x sample expression dataframe
  381. expression = utils.check_dict(expression)
  382. probe_exp = pd.concat(expression.values(), axis=1)
  383. # determine aggregation function based on provided method; also reduce
  384. # probe expression if required (i.e., max_variance, max_intensity)
  385. if method == 'max_variance':
  386. probe_exp = pd.DataFrame(probe_exp.std(axis=1), columns=[method])
  387. agg = functools.partial(_max_idx, column=method)
  388. elif method == 'max_intensity':
  389. probe_exp = pd.DataFrame(probe_exp.mean(axis=1), columns=[method])
  390. # probe_exp.name = method
  391. agg = functools.partial(_max_idx, column=method)
  392. elif method == 'pc_loading':
  393. agg = _max_loading
  394. elif method in ['corr_variance', 'corr_intensity']:
  395. agg = functools.partial(_correlate, method=method[5:])
  396. else:
  397. raise ValueError(f'Provided method {method} is invalid. Please check '
  398. 'inputs and try again.')
  399. info = pd.merge(probes[['gene_symbol']], probe_exp, on='probe_id')
  400. return _groupby_and_apply(expression, probes, info, agg)
  401. _max_variance = functools.partial(_collapse, method='max_variance')
  402. _max_intensity = functools.partial(_collapse, method='max_intensity')
  403. _pc_loading = functools.partial(_collapse, method='pc_loading')
  404. _corr_variance = functools.partial(_collapse, method='corr_variance')
  405. _corr_intensity = functools.partial(_collapse, method='corr_intensity')
  406. SELECTION_METHODS = dict(
  407. mean=_average,
  408. average=_average,
  409. max_intensity=_max_intensity,
  410. max_variance=_max_variance,
  411. pc_loading=_pc_loading,
  412. corr_variance=_corr_variance,
  413. corr_intensity=_corr_intensity,
  414. diff_stability=_diff_stability,
  415. rnaseq=_rnaseq
  416. )
  417. AGG_METHODS = [ # can only be used with `donor_probes='aggregate'`
  418. 'mean', 'average', 'diff_stability', 'rnaseq'
  419. ]
  420. COLLAPSE_METHODS = [ # methods that don't SELECT but COLLAPSE ACROSS probes
  421. 'mean', 'average'
  422. ]
  423. def collapse_probes(microarray, annotation, probes, method='diff_stability',
  424. donor_probes='aggregate'):
  425. """
  426. Reduces `microarray` to a sample x gene expression dataframe
  427. Using provided `method`, reduces `microarray` expression data by either (1)
  428. selecting a representative probe amongst all probes indexing the same gene,
  429. or (2) collapsing across all probes indexing the same gene. See Notes for
  430. more information on different methods available.
  431. Parameters
  432. ----------
  433. microarray : dict of str or pandas.DataFrame
  434. Dictionary where keys are donor IDs and values are filepaths to (or
  435. dataframes of) MicroarrayExpression.csv files from Allen Brain
  436. Institute
  437. annotation : dict of str or pandas.DataFrame
  438. Dictionary where keys are donor IDs and values are filepaths to (or
  439. dataframes of) SampleAnnot.csv files from Allen Brain Institute. Only
  440. used if `method='diff_stability'`
  441. probes : str or pandas.DataFrame
  442. Filepath to Probes.csv or dataframe containing information on
  443. microarray probes that should be considered in representative analysis.
  444. Generally intensity-based filtering (i.e., via :func:`filter_probes()`)
  445. should have been used to reduce this list to only those probes with
  446. good expression signal
  447. method : str, optional
  448. Selection method for subsetting (or collapsing across) probes from the
  449. same gene. Must be one of 'average', 'max_intensity', 'max_variance',
  450. 'pc_loading', 'corr_intensity', 'corr_variance', 'diff_stability', or
  451. 'rnaseq'; see Notes for more information. Default: 'diff_stability'
  452. donor_probes : str, optional
  453. Whether specified `probe_selection` method should be performed with
  454. microarray data from all donors ('aggregate'), independently for each
  455. donor ('independent'), or based on the most common selected probe
  456. across donors ('common'). Not all combinations of `probe_selection`
  457. and `donor_probes` methods are viable. Default: 'aggregate'
  458. Returns
  459. -------
  460. expression : dict of (S, G) pandas.DataFrame
  461. Dictionary where keys are donor IDs and values are dataframes with `S`
  462. rows representing distinct samples and `G` columns representing unique
  463. genes. Entries of dataframe indicate microarray expression levels for
  464. each combination of sample + gene. Columns will be identical across all
  465. dataframes, but `S` will vary by donor.
  466. Notes
  467. -----
  468. The following methods can be used for collapsing across probes when
  469. multiple probes are available for the same gene.
  470. 1. ``method='average'``
  471. Uses the average of expression data across probes indexing the same gene as
  472. in [PR3]_, [PR5]_, [PR9]_, [PR10]_, [PR15]_, [PR16]_, and [PR17]_. Using
  473. `method='mean'` will do the same thing.
  474. 2. ``method='max_intensity'``
  475. Selects probe with maximum average expression across samples from all
  476. donors as in [PR14]_.
  477. 3. ``method='max_variance'``
  478. Selects probe with maximum variance in expression across samples from all
  479. donors as in [PR12]_.
  480. 4. ``method='pc_loading'``
  481. Selects probe with maximum loading on first principal component of
  482. decomposition performed across samples from all donors as in [PR13]_.
  483. 5. ``method='corr_intensity'``
  484. Selects probe with maximum correlation to other probes from same gene when
  485. >2 probes exist; otherwise, uses same procedure as `method=max_intensity`.
  486. Used in [PR1]_ and [PR11]_.
  487. 6. ``method='corr_variance'``
  488. Selects probe with maximum correlation to other probes from same gene when
  489. >2 probes exist; otherwise, uses same procedure as `method=max_variance`.
  490. Used in [PR2]_, [PR4]_, and [PR6]_.
  491. 7. ``method='diff_stability'``
  492. Selects probe with the most consistent pattern of regional variation across
  493. donors (i.e., highest average correlation across brain regions between all
  494. pairs of donors) as in [PR7]_ and [PR8]_.
  495. 8. ``method='rnaseq'``
  496. Selects probes with most consistent pattern of regional variation to RNAseq
  497. data (across the two donors with RNAseq data).
  498. References
  499. ----------
  500. .. [PR1] Anderson, K. M., Krienen, F. M., Choi, E. Y., Reinen, J. M., Yeo,
  501. B. T., & Holmes, A. J. (2018). Gene expression links functional networks
  502. across cortex and striatum. Nature Communications, 9(1), 1428.
  503. .. [PR2] Burt, J. B., Demirtaş, M., Eckner, W. J., Navejar, N. M., Ji, J.
  504. L., Martin, W. J., ... & Murray, J. D. (2018). Hierarchy of
  505. transcriptomic specialization across human cortex captured by structural
  506. neuroimaging topography. Nature Neuroscience, 21(9), 1251.
  507. .. [PR3] Eising, E., Huisman, S. M., Mahfouz, A., Vijfhuizen, L. S.,
  508. Anttila, V., Winsvold, B. S., ... & Boomsma, D. I. (2016). Gene
  509. co-expression analysis identifies brain regions and cell types involved
  510. in migraine pathophysiology: a GWAS-based study using the Allen Human
  511. Brain Atlas. Human Genetics, 135(4), 425-439.
  512. .. [PR4] Forest, M., Iturria‐Medina, Y., Goldman, J. S., Kleinman, C. L.,
  513. Lovato, A., Oros Klein, K., ... & Greenwood, C. M. (2017). Gene networks
  514. show associations with seed region connectivity. Human Brain Mapping,
  515. 38(6), 3126-3140.
  516. .. [PR5] French, L., & Paus, T. (2015). A FreeSurfer view of the cortical
  517. transcriptome generated from the Allen Human Brain Atlas. Frontiers in
  518. Neuroscience, 9, 323.
  519. .. [PR6] Hawrylycz, M. J., Lein, E. S., Guillozet-Bongaarts, A. L., Shen,
  520. E. H., Ng, L., Miller, J. A., ... & Abajian, C. (2012). An anatomically
  521. comprehensive atlas of the adult human brain transcriptome. Nature,
  522. 489(7416), 391.
  523. .. [PR7] Hawrylycz, M., Miller, J. A., Menon, V., Feng, D., Dolbeare, T.,
  524. Guillozet-Bongaarts, A. L., ... & Glasser, M. F. (2015). Canonical
  525. genetic signatures of the adult human brain. Nature Neuroscience,
  526. 18(12), 1832.
  527. .. [PR8] Kirsch, L., & Chechik, G. (2016). On expression patterns and
  528. developmental origin of human brain regions. PLoS Computational Biology,
  529. 12(8), e1005064.
  530. .. [PR9] Krienen, F. M., Yeo, B. T., Ge, T., Buckner, R. L., & Sherwood, C.
  531. C. (2016). Transcriptional profiles of supragranular-enriched genes
  532. associate with corticocortical network architecture in the human brain.
  533. Proceedings of the National Academy of Sciences, 113(4), E469-E478.
  534. .. [PR10] McColgan, P., Gregory, S., Seunarine, K. K., Razi, A., Papoutsi,
  535. M., Johnson, E., ... & Scahill, R. I. (2018). Brain regions showing
  536. white matter loss in Huntington’s disease are enriched for synaptic and
  537. metabolic genes. Biological Psychiatry, 83(5), 456-465.
  538. .. [PR11] Myers, E. M., Bartlett, C. W., Machiraju, R., & Bohland, J. W.
  539. (2015). An integrative analysis of regional gene expression profiles in
  540. the human brain. Methods, 73, 54-70.S
  541. .. [PR12] Negi, S. K., & Guda, C. (2017). Global gene expression profiling
  542. of healthy human brain and its application in studying neurological
  543. disorders. Scientific Reports, 7(1), 897.
  544. .. [PR13] Parkes, L., Fulcher, B. D., Yücel, M., & Fornito, A. (2017).
  545. Transcriptional signatures of connectomic subregions of the human
  546. striatum. Genes, Brain and Behavior, 16(7), 647-663.
  547. .. [PR14] Romero-Garcia, R., Whitaker, K. J., Váša, F., Seidlitz, J.,
  548. Shinn, M., Fonagy, P., ... & Vértes, P. E. (2018). Structural covariance
  549. networks are coupled to expression of genes enriched in supragranular
  550. layers of the human cortex. NeuroImage, 171, 256-267.
  551. .. [PR15] Tan, P. P. C., French, L., & Pavlidis, P. (2013). Neuron-enriched
  552. gene expression patterns are regionally anti-correlated with
  553. oligodendrocyte-enriched patterns in the adult mouse and human brain.
  554. Frontiers in Neuroscience, 7, 5.
  555. .. [PR16] Vértes, P. E., Rittman, T., Whitaker, K. J., Romero-Garcia, R.,
  556. Váša, F., Kitzbichler, M. G., ... & Goodyer, I. M. (2016). Gene
  557. transcription profiles associated with inter-modular hubs and connection
  558. distance in human functional magnetic resonance imaging networks.
  559. Philosophical Transactions of the Royal Society B: Biological Sciences,
  560. 371(1705), 20150362.
  561. .. [PR17] Whitaker, K. J., Vértes, P. E., Romero-Garcia, R., Váša, F.,
  562. Moutoussis, M., Prabhu, G., ... & Tait, R. (2016). Adolescence is
  563. associated with genomically patterned consolidation of the hubs of the
  564. human brain connectome. Proceedings of the National Academy of Sciences,
  565. 113(32), 9105-9110.
  566. """
  567. try:
  568. collfunc = SELECTION_METHODS[method]
  569. except KeyError:
  570. raise ValueError(f'Provided `method` "{method}" is invalid; must be '
  571. f'one of {list(SELECTION_METHODS)}')
  572. valid_probes = ['aggregate', 'common', 'independent']
  573. if donor_probes not in valid_probes:
  574. raise ValueError(f'Provided `donor_probes` "{donor_probes}" is '
  575. f'invalid; must be one of {valid_probes}')
  576. LGR.info(f'Reducing probes indexing same gene with method: {method}')
  577. # subset microarray data for pre-selected probes + samples
  578. # this will also left/right mirror samples, if previously requested
  579. probes = io.read_probes(probes)
  580. microarray = utils.check_dict(microarray)
  581. annotation = utils.check_dict(annotation)
  582. for donor, micro in microarray.items():
  583. samp = io.read_annotation(annotation[donor]).index
  584. microarray[donor] = io.read_microarray(micro).loc[probes.index, samp]
  585. # now, "collect" the probes based on the provided `method`
  586. if method in AGG_METHODS or donor_probes == 'aggregate':
  587. # perform the collection function for all donors, together
  588. microarray = collfunc(microarray, probes, annotation)
  589. elif donor_probes == 'independent':
  590. # perform the collection function for each donor separately
  591. for donor in microarray:
  592. microarray.update(collfunc({donor: microarray[donor]}, probes,
  593. {donor: annotation[donor]}))
  594. elif donor_probes == 'common':
  595. # perform collection function for each donor separately and retain ONLY
  596. # the chose probe IDs
  597. probe_ids = [
  598. collfunc(
  599. {donor: microarray[donor]}, probes, {donor: annotation[donor]}
  600. )[donor].columns
  601. for donor in microarray
  602. ]
  603. # find the mode of the probe IDs chosen across donors and then subset
  604. # those probes from the original microarray dataframes for all donors
  605. probe_ids = np.squeeze(sstats.mode(probe_ids, axis=0)[0])
  606. for donor in microarray:
  607. microarray[donor] = microarray[donor].loc[probe_ids].T
  608. # convert probe IDs as column names to gene symbols
  609. if method not in COLLAPSE_METHODS:
  610. for donor, micro in microarray.items():
  611. symbols = probes.loc[micro.columns, 'gene_symbol']
  612. micro = micro.set_axis(symbols, axis=1, copy=True)
  613. microarray[donor] = micro.sort_index(axis=1)
  614. n_genes = utils.first_entry(microarray).shape[-1]
  615. LGR.info(f'{n_genes} genes remain after probe filtering + selection')
  616. return microarray

probes_.py at commit dc4a007, under BSD-3-Clause · at the source

Overview

Authors: Meng-Xuan Qiao1, Wei Wei1,2,3,4, Meng Zhou1, Ming-li Li5, Ya-min Zhang1,2,3,4, Xiao-Jing Li1,2,3,4, Wei Deng1,2,3,4, Wan-Jun Guo1,2,3,4, Qiang Wang5, Hua Yu1,2,3,4, Tao Li1,2,3,4
  1. Affiliated Mental Health Center & Hangzhou Seventh People’s Hospital, School of Brain Science and Brain Medicine, Zhejiang University School of Medicine, Hangzhou, Zhejiang 310058 China
  2. Zhejiang Key Laboratory of Clinical and Basic Research on Mental Disorders, Hangzhou, Zhejiang China
  3. Zhejiang Engineering Research Center for Intelligent Diagnosis and Treatment of Mental Disorders, Hangzhou, Zhejiang China
  4. Liangzhu Laboratory, MOE Frontier Science Center for Brain Science and Brain- machine Integration, State Key Laboratory of Brain-machine Intelligence, Zhejiang University, Hangzhou, Zhejiang China
  5. Mental Health Center, West China Hospital of Sichuang University, Chengdu, Sichuan China
Journal: BMC psychiatry, volume 26, issue 1, article 600
Dates: received 4 March 2026; accepted 19 May 2026; published online 30 May 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1186/s12888-026-08216-5 · PMID 42218423 · PMCID PMC13445930 · OpenAlex W7162855681
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), fMRI (modality), human (organism), bipolar (population), cellular / molecular (subfield)
Methods: Spectral & time-frequency, Connectivity, Statistics, Smoothing, state filtering, decompositions, Preprocessing, fMRI & imaging
Keywords: Bipolar disorder, Habenular volume, Resting-state functional connectivity, Biological mechanisms, Polygenic risk score
MeSH: Bipolar Disorder*, Habenula*, Transcriptome*, Adult, Female, Genetic Risk Score, Humans, Magnetic Resonance Imaging, Male, Middle Aged, Multimodal Imaging, Neural Pathways (* major topic)
Topic: Bipolar Disorder and Treatment (Psychiatry and Mental health, Medicine), according to OpenAlex
Funding: Construction Fund of Key Medical Disciplines of Hangzhou (2025HZGF10); National Natural Science Foundation of China (82101598, 82230046); China Brain Project (STI2030-2021ZD0200404); Zhejiang Clinovation Pride (CXTD202501053)
Citations: not cited yet (Europe PMC); 54 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.

rmarkello/abagen

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: dc4a007e4e902e51f97251390c8d1bbf7e58c6d3, 29 September 2023
Languages: Python (52), Shell (1)
Size: 130 files, 53 scripts
Software Heritage: archived
Found in: the text, “Spatial transcriptome in the brain”
Holds: README, license file, environment (requirements.txt, setup.cfg, setup.py, docs/requirements.txt), tests, continuous integration, documentation
Not found: CITATION.cff
Tools: NumPy (23 files), pandas (22 files), abagen (20 files), NiBabel (11 files), SciPy (7 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
55 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 53 scripts, each with its path and the digest of its content;
  • 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it says that the data are available on request

Read it in the paper: doi.org/10.1186/s12888-026-08216-5.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 11 authors, 5 keywords, 12 MeSH terms, 4 funders, 52 references.

Cite

This paper

Qiao, M.-X., Wei, W., Zhou, M., Li, M.-l., Zhang, Y.-m., Li, X.-J., Deng, W., Guo, W.-J., Wang, Q., Yu, H., & Li, T. (2026). Habenular structural-functional dysconnectivity in bipolar disorder: evidence from multimodal imaging and transcriptomic integration. BMC psychiatry, 26(1), 600. https://doi.org/10.1186/s12888-026-08216-5

BibTeX

@article{qiao2026habenular,
author = {Qiao, Meng-Xuan and Wei, Wei and Zhou, Meng and Li, Ming-li and Zhang, Ya-min and Li, Xiao-Jing and Deng, Wei and Guo, Wan-Jun and Wang, Qiang and Yu, Hua and Li, Tao},
title = {{Habenular structural-functional dysconnectivity in bipolar disorder: evidence from multimodal imaging and transcriptomic integration}},
journal = {BMC psychiatry},
year = {2026},
month = may,
volume = {26},
number = {1},
pages = {600},
publisher = {BMC},
issn = {1471-244X},
doi = {10.1186/s12888-026-08216-5},
url = {https://doi.org/10.1186/s12888-026-08216-5},
pmid = {42218423},
pmcid = {PMC13445930}
}

RIS

TY - JOUR
AU - Qiao, Meng-Xuan
AU - Wei, Wei
AU - Zhou, Meng
AU - Li, Ming-li
AU - Zhang, Ya-min
AU - Li, Xiao-Jing
AU - Deng, Wei
AU - Guo, Wan-Jun
AU - Wang, Qiang
AU - Yu, Hua
AU - Li, Tao
TI - Habenular structural-functional dysconnectivity in bipolar disorder: evidence from multimodal imaging and transcriptomic integration
T2 - BMC psychiatry
J2 - BMC Psychiatry
PY - 2026
DA - 2026/05/30
VL - 26
IS - 1
SP - 600
SN - 1471-244X
PB - BMC
DO - 10.1186/s12888-026-08216-5
UR - https://doi.org/10.1186/s12888-026-08216-5
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s12888-026-08216-5",
"type": "article-journal",
"title": "Habenular structural-functional dysconnectivity in bipolar disorder: evidence from multimodal imaging and transcriptomic integration",
"container-title": "BMC psychiatry",
"author": [
{
"family": "Qiao",
"given": "Meng-Xuan"
},
{
"family": "Wei",
"given": "Wei"
},
{
"family": "Zhou",
"given": "Meng"
},
{
"family": "Li",
"given": "Ming-li"
},
{
"family": "Zhang",
"given": "Ya-min"
},
{
"family": "Li",
"given": "Xiao-Jing"
},
{
"family": "Deng",
"given": "Wei"
},
{
"family": "Guo",
"given": "Wan-Jun"
},
{
"family": "Wang",
"given": "Qiang"
},
{
"family": "Yu",
"given": "Hua"
},
{
"family": "Li",
"given": "Tao"
}
],
"container-title-short": "BMC Psychiatry",
"volume": "26",
"issue": "1",
"page": "600",
"DOI": "10.1186/s12888-026-08216-5",
"PMID": "42218423",
"PMCID": "PMC13445930",
"ISSN": "1471-244X",
"publisher": "BMC",
"URL": "https://doi.org/10.1186/s12888-026-08216-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
30
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-74153-2 [code]
Regional, functional and transcriptomic decoding of multidimensional brain structure alterations in obsessive-compulsive disorder.
Journal: Nature communications
In common: abagen, NiBabel, pandas, 2 other tools, genetics / omics, cellular / molecular, 1 reference
[2] doi:10.1017/s0033291726104838 [code]
Variations of structural-functional coupling in post-traumatic stress disorder are associated with underlying molecular and transcriptional features.
Journal: Psychological medicine
In common: abagen, NiBabel, pandas, 2 other tools, cellular / molecular, 1 reference
[3] doi:10.1038/s42003-025-09444-3 [code]
Decoupling of neurophysiological activity from structure mirrors global microarchitectural and neuromodulatory trends.
Journal: Communications biology
In common: abagen, NiBabel, pandas, 2 other tools, cellular / molecular, 1 reference
[4] doi:10.3389/fnins.2026.1833695 [code]
Transcriptional signatures of the cortical morphometric similarity network gradient in left temporal lobe epilepsy with different seizure symptoms.
Journal: Frontiers in neuroscience
In common: abagen, NiBabel, pandas, 2 other tools, 1 reference
[5] doi:10.1038/s41467-026-71719-y [code]
Brain functional-structural gradient coupling reflects development, behavior and genetic influences.
Journal: Nature communications
In common: abagen, NiBabel, pandas, 2 other tools, 1 reference
[6] doi:10.1038/s42003-026-09956-6 [code]
Linking changes in sulcal morphometry to cognitive development from childhood to adolescence.
Journal: Communications biology
In common: abagen, NiBabel, pandas, 2 other tools, 1 reference
[7] doi:10.1162/imag.a.1342 [code]
Hyperpolarized &lt;sup&gt;13&lt;/sup&gt;C-bicarbonate production correlates with excitatory neuron distribution in the human brain.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: abagen, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 1 reference
[8] doi:10.1038/s43856-026-01707-2 [code]
Decreased amyloid-related structure-function coupling in preclinical Alzheimer's disease.
Journal: Communications medicine
In common: abagen, NiBabel, pandas, 2 other tools
[9] doi:10.1371/journal.pmed.1004809 [code]
Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data.
Journal: PLoS medicine
In common: abagen, NiBabel, pandas, 2 other tools
[10] doi:10.1038/s41467-026-71568-9 [code]
Convergent and selective representations of pain, appetitive processes, aversive processes, and cognitive control in the insula.
Journal: Nature communications
In common: abagen, NiBabel, pandas, 2 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.