OSCR

Learning brain dynamics across distinct scaling regimes reveals psychiatric signatures.

Code ↔ Paper

22 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 22 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Methods › Experimental settings ↔ visualization.py, lines 24–108 · score 0.91 · Gradient clipping, AdamW, hidden layers, weight decay, model architecture, accumulation
  2. [2] § Methods › Experimental settings ↔ main.py, lines 47–95 · score 0.83 · Gradient clipping, AdamW, weight decay, accumulation, norm, SGDR
  3. [3] § Methods › Statistics and reproducibility › Statistical analysis ↔ loss_writer.py, lines 79–136 · score 0.81 · balanced accuracy, F1 score, binary classification, NMSE, R2, sensitivity
  4. [4] § Methods › Frequency-resolved decomposition based on scale-free principles › Step 1: Ultralow-frequency boundary identification using Lorentzian fitting ↔ data_preprocess_and_load/datasets.py, lines 37–78 · score 0.67 · curve_fit, power spectrum, Lorentzian, Ultralow
  5. [5] § Methods › Communicability-based pretraining strategy › Implementation details and dataset ↔ data_preprocess_and_load/datasets.py, lines 239–320 · score 0.65 · UK Biobank, downstream tasks, sequence length, fine tuning, Pretraining, band
  6. [6] § Results › Experimental setup and downstream tasks › MBBN demonstrates superior performance across diverse neuroimaging tasks ↔ data_preprocess_and_load/datasets.py, lines 239–320 · score 0.64 · fluid intelligence, downstream tasks, HCP MMP1, phenotypes, Schaefer, UKB
  7. [7] § Methods › Communicability-based pretraining strategy › Masking strategy and pretraining loss ↔ model.py, lines 296–379 · score 0.63 · MBBN pretraining, temporal masking, spatial masking, windows, zeros, head
  8. [8] § Methods › Statistics and reproducibility › Statistical analysis ↔ metrics.py, the whole file · a weak match · score 0.63 · F1 score, NMSE, R2, sensitivity, metrics, MAE
  9. [9] § Methods › Preprocessing ↔ data_preprocess_and_load/datasets.py, lines 325–424 · score 0.62 · Adolescent Brain Cognitive, sequence length, ABCD, Preprocessing
  10. [10] § Methods › Communicability-based pretraining strategy › Masking strategy and pretraining loss ↔ main.py, lines 97–164 · score 0.61 · MBBN pretraining, temporal masking, spatial masking, windows, nodes, head
  11. [11] § Methods › Statistics and reproducibility › Data splitting and sample sizes › Reproducibility ↔ environment.sh, the whole file · a weak match · score 0.61 · nibabel, conda, nitime, scikit, PyTorch, Python
  12. [12] § Methods › Communicability-based pretraining strategy › Masking strategy and pretraining loss ↔ model.py, lines 296–379 · score 0.58 · Temporal masking, Spatial masking, windows, zero, ROIs, Pretraining
  13. [13] § Methods › Preprocessing ↔ data_preprocess_and_load/ROI_EXTRACT_UKB.py, lines 32–40 · score 0.58 · MNI space, HCP MMP1, ROI, preprocessing
  14. [14] § Methods › Frequency-resolved decomposition based on scale-free principles › Validation of frequency-specific scaling properties ↔ data_preprocess_and_load/datasets.py, lines 37–78 · score 0.57 · spline multifractal, knee frequencies, spectrum, Lorentzian, power, fit
  15. [15] § Results › Experimental setup and downstream tasks › MBBN demonstrates superior performance across diverse neuroimaging tasks ↔ main.py, lines 7–45 · score 0.55 · vanilla BERT, model architectures, fine tuning, scratch, baseline, seed
  16. [16] § Methods › Model design › Model architecture ↔ visualization.py, lines 24–108 · score 0.55 · sequence length, spatial loss, head, UKB, ABIDE, architecture
  17. [17] § Methods › Dataset description and subject selection criteria ↔ data_preprocess_and_load/datasets.py, lines 325–424 · score 0.54 · Adolescent Brain Cognitive, MRI, ROI, preprocessing, ABCD
  18. [18] § Methods › Frequency-resolved decomposition based on scale-free principles › Validation of frequency-specific scaling properties ↔ data_preprocess_and_load/datasets.py, lines 133–234 · score 0.54 · knee frequencies, 1.5 s, bounded, TR, ABIDE, fMRI
  19. [19] § Methods › Communicability-based pretraining strategy › Masking strategy and pretraining loss ↔ main.py, lines 97–164 · score 0.54 · Temporal masking, Spatial masking, windows, nodes, Pretraining, loss
  20. [20] § Methods › Statistics and reproducibility › Data splitting and sample sizes › Reproducibility ↔ utils.py, lines 112–122 · score 0.52 · random module, reproducibility, Python, PyTorch, CUDA, seeds
  21. [21] § Methods › Statistics and reproducibility › Data splitting and sample sizes ↔ data_preprocess_and_load/datasets.py, lines 81–110 · score 0.52 · zero variance ROIs, split
  22. [22] § Methods › Model design › Model architecture ↔ main.py, lines 47–95 · score 0.51 · sequence length, spatial loss, UKB, ABIDE, ABCD, MBBN

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 424 lines · 16 KB · MIT · 8 matches

  1. import numpy as np
  2. import pandas as pd
  3. import os
  4. import torch
  5. from torch.utils.data import Dataset
  6. import torch.nn.functional as F
  7. import scipy
  8. from scipy import stats
  9. from scipy.optimize import curve_fit
  10. from iminuit import Minuit
  11. from numba import njit
  12. import warnings
  13. warnings.filterwarnings("ignore")
  14. from nitime.timeseries import TimeSeries
  15. from nitime.analysis import SpectralAnalyzer, FilterAnalyzer
  16. @njit
  17. def lorentzian_function(x, s0, f1):
  18. """Lorentzian power spectral density model (Eq. 2 in paper)."""
  19. return (s0 * f1**2) / (x**2 + f1**2)
  20. @njit
  21. def spline_multifractal(x, beta_low, beta_high, A, f2, smoothness):
  22. """Spline multifractal model with cubic-spline transition (Eq. 3-4 in paper)."""
  23. log_x = np.log(x)
  24. log_f2 = np.log(f2)
  25. w = np.where(log_x < log_f2 - smoothness, 0,
  26. np.where(log_x > log_f2 + smoothness, 1,
  27. 0.5 * (1 - np.cos(np.pi * (log_x - log_f2 + smoothness) / (2 * smoothness)))))
  28. return A * x**(beta_low * (1 - w) + beta_high * w)
  29. def _find_knee_frequencies(y, TR, seq_len, intermediate_vec,
  30. lz_p0, lz_bounds):
  31. """Identify f1 (ultralow/low) and f2 (low/high) knee frequencies.
  32. Step 1: Lorentzian fit on the full PSD → f1
  33. Step 2: Spline multifractal fit on PSD above f1 → f2
  34. Returns (f1, f2) in Hz.
  35. """
  36. # Average over ROIs to get a single representative power spectrum
  37. sample_whole = y.mean(axis=0)
  38. T = TimeSeries(sample_whole, sampling_interval=TR)
  39. S = SpectralAnalyzer(T)
  40. xdata = np.array(S.spectrum_fourier[0][1:])
  41. ydata = np.abs(S.spectrum_fourier[1][1:])
  42. # Step 1 — Lorentzian: find f1
  43. popt, _ = curve_fit(lorentzian_function, xdata, ydata,
  44. p0=lz_p0, bounds=lz_bounds, maxfev=50000)
  45. f1 = popt[-1]
  46. knee = round(f1 / (1 / (seq_len * TR)))
  47. if knee <= 0:
  48. knee = 1
  49. # Step 2 — Spline multifractal: find f2
  50. def least_squares(beta_low, beta_high, A, f2, smoothness):
  51. y_pred = spline_multifractal(xdata[knee:], beta_low, beta_high,
  52. A, f2, smoothness)
  53. return np.sum((y_pred - ydata[knee:])**2)
  54. m = Minuit(least_squares, beta_low=-1.2, beta_high=-0.5,
  55. A=10, f2=0.08, smoothness=0.25)
  56. m.limits['beta_low'] = (-5, -0.01)
  57. m.limits['beta_high'] = (-5, -0.01)
  58. m.limits['f2'] = (f1 + 0.0001, 0.15)
  59. m.limits['A'] = (1, 30)
  60. m.limits['smoothness'] = (0.001, 1)
  61. m.migrad()
  62. f2 = m.values['f2']
  63. return f1, f2
  64. def _filter_three_bands(y, TR, f1, f2, filtering_type):
  65. """Split timeseries y [ROI × time] into ultralow, low, and high bands.
  66. Uses f2 as the high-pass cutoff and f1 as the low-pass cutoff.
  67. FIR is used for ultralow/low; Boxcar for high (as in Table 7).
  68. """
  69. # High frequency (above f2)
  70. T1 = TimeSeries(y, sampling_interval=TR)
  71. FA1 = FilterAnalyzer(T1, lb=f2)
  72. if filtering_type == 'FIR':
  73. high = FA1.fir.data
  74. # Guard against zero-variance ROIs after FIR filtering
  75. std = np.std(high, axis=1, keepdims=True)
  76. high = (high - high.mean(axis=1, keepdims=True)) / (std + 1e-10)
  77. ultralow_low = FA1.data - FA1.fir.data
  78. else: # Boxcar
  79. high = stats.zscore(FA1.filtered_boxcar.data, axis=1)
  80. ultralow_low = FA1.data - FA1.filtered_boxcar.data
  81. # Low and ultralow frequencies (split at f1)
  82. T2 = TimeSeries(ultralow_low, sampling_interval=TR)
  83. FA2 = FilterAnalyzer(T2, lb=f1)
  84. if filtering_type == 'FIR':
  85. low = stats.zscore(FA2.fir.data, axis=1)
  86. ultralow = stats.zscore(FA2.data - FA2.fir.data, axis=1)
  87. else: # Boxcar
  88. low = stats.zscore(FA2.filtered_boxcar.data, axis=1)
  89. ultralow = stats.zscore(FA2.data - FA2.filtered_boxcar.data, axis=1)
  90. return ultralow, low, high
  91. class BaseDataset(Dataset):
  92. def __init__(self):
  93. super().__init__()
  94. def register_args(self, **kwargs):
  95. self.index_l = []
  96. self.target = kwargs.get('target')
  97. self.fine_tune_task = kwargs.get('fine_tune_task')
  98. self.dataset_name = kwargs.get('dataset_name')
  99. self.intermediate_vec = kwargs.get('intermediate_vec')
  100. self.filtering_type = kwargs.get('filtering_type')
  101. self.sequence_length = kwargs.get('sequence_length')
  102. self.pretrained_model_weights_path = kwargs.get('pretrained_model_weights_path')
  103. self.finetune = kwargs.get('finetune')
  104. self.transfer_learning = bool(self.pretrained_model_weights_path) or self.finetune
  105. self.finetune_test = kwargs.get('finetune_test')
  106. # ─── ABIDE ────────────────────────────────────────────────────────────────────
  107. class ABIDE_fMRI_timeseries(BaseDataset):
  108. """ABIDE-I/II dataset. Target: ASD (binary classification)."""
  109. # TR values per acquisition site (in seconds)
  110. _SITE_TR = {
  111. 'CALTECH': 2.0, 'CMU': 2.0, 'KKI': 2.5, 'LEUVEN': 1.66,
  112. 'MAX_MUN': 3.0, 'NYU': 2.0, 'OHSU': 2.5, 'OLIN': 1.5,
  113. 'PITT': 1.5, 'SBL': 2.2, 'SDSU': 2.0, 'STANFORD': 2.0,
  114. 'TRINITY': 2.0, 'UCLA': 3.0, 'UM': 2.0, 'USM': 2.0, 'YALE': 2.0,
  115. }
  116. def __init__(self, **kwargs):
  117. self.register_args(**kwargs)
  118. self.data_dir = kwargs.get('abide_path')
  119. self.meta_data = pd.read_csv(
  120. os.path.join(kwargs.get('base_path'), 'data', 'metadata',
  121. 'ABIDE1+2_meta.csv'))
  122. self.site_meta_data = pd.read_csv(
  123. os.path.join(kwargs.get('base_path'), 'data', 'metadata',
  124. 'ABIDE1_pheno_and_sites.csv'))
  125. # Build file list
  126. data_list = []
  127. for sub in os.listdir(self.data_dir):
  128. if self.intermediate_vec == 400:
  129. data_list.append(os.path.join(
  130. self.data_dir, sub,
  131. 'schaefer_400Parcels_17Networks_' + sub + '.npy'))
  132. elif sub.startswith('00'): # ABIDE-I subjects
  133. data_list.append(os.path.join(
  134. self.data_dir, sub, 'hcp_mmp1_360_' + sub + '.npy'))
  135. if self.intermediate_vec == 360:
  136. unified_name_list = [i[2:] for i in os.listdir(self.data_dir)
  137. if i.startswith('00')]
  138. else:
  139. unified_name_list = [i[2:] if i.startswith('00') else i
  140. for i in os.listdir(self.data_dir)]
  141. non_na = self.meta_data.dropna(axis=0)
  142. valid_sub = (set(str(i) for i in non_na['SUB_ID'])
  143. & set(unified_name_list))
  144. for filename in data_list:
  145. sub = filename.split('/')[-2]
  146. if sub.startswith('00'): # ABIDE-I
  147. subid = sub[2:]
  148. site_row = self.site_meta_data[
  149. self.site_meta_data['SUB_ID'] == int(subid)]['SITE_ID']
  150. site = site_row.values[0] if len(site_row) else 'ABIDE2'
  151. else:
  152. subid = sub
  153. site = 'ABIDE2'
  154. if subid in valid_sub:
  155. if self.target == 'sex':
  156. target_val = non_na.loc[
  157. non_na['SUB_ID'] == int(subid), 'SEX'].values[0]
  158. else: # ASD
  159. target_val = non_na.loc[
  160. non_na['SUB_ID'] == int(subid), 'DX_GROUP'].values[0]
  161. target = torch.tensor(1.0 if target_val == 2 else 0.0)
  162. self.index_l.append((subid, sub, filename, target, site))
  163. def __len__(self):
  164. return len(self.index_l)
  165. def __getitem__(self, index):
  166. subj, subj_name, path_to_fMRIs, target, site = self.index_l[index]
  167. y = np.load(path_to_fMRIs)[:self.sequence_length].T # [ROI, seq_len]
  168. # Padding to pretrained sequence length (464) for finetuning
  169. pad = 464 - self.sequence_length
  170. # Site-specific TR
  171. TR = next((v for k, v in self._SITE_TR.items() if k in site), 3.0)
  172. f1, f2 = _find_knee_frequencies(
  173. y, TR, self.sequence_length, self.intermediate_vec,
  174. lz_p0=[900, 0.05],
  175. lz_bounds=([0, 0.01], [1200, 0.1]))
  176. ultralow, low, high = _filter_three_bands(
  177. y, TR, f1, f2, self.filtering_type)
  178. # Always pad to match pretraining length (464)
  179. high = F.pad(torch.from_numpy(high),
  180. (pad // 2, pad // 2), 'constant', 0).T.float()
  181. low = F.pad(torch.from_numpy(low),
  182. (pad // 2, pad // 2), 'constant', 0).T.float()
  183. ultralow = F.pad(torch.from_numpy(ultralow),
  184. (pad // 2, pad // 2), 'constant', 0).T.float()
  185. return {
  186. 'fmri_highfreq_sequence': high,
  187. 'fmri_lowfreq_sequence': low,
  188. 'fmri_ultralowfreq_sequence': ultralow,
  189. 'subject': subj,
  190. 'subject_name': subj_name,
  191. self.target: target,
  192. }
  193. # ─── UKB ──────────────────────────────────────────────────────────────────────
  194. class UKB_fMRI_timeseries(BaseDataset):
  195. """UK Biobank dataset. Used for pretraining (target='reconstruction')
  196. and downstream tasks (sex, depression, fluid intelligence)."""
  197. TR = 0.735 # seconds
  198. def __init__(self, **kwargs):
  199. self.register_args(**kwargs)
  200. self.data_dir = kwargs.get('ukb_path')
  201. self.meta_data = pd.read_csv(
  202. os.path.join(kwargs.get('base_path'), 'data', 'metadata',
  203. 'UKB_phenotype_gps_fluidint.csv'))
  204. valid_sub = list(map(int, os.listdir(self.data_dir)))
  205. if self.target != 'reconstruction':
  206. non_na = self.meta_data[['eid', self.target]].dropna(axis=0)
  207. subjects = list(set(non_na['eid']) & set(valid_sub))
  208. else:
  209. subjects = valid_sub
  210. if self.fine_tune_task == 'regression':
  211. cont_mean = non_na[self.target].mean()
  212. cont_std = non_na[self.target].std()
  213. self.mean = cont_mean
  214. self.std = cont_std
  215. for i, subject in enumerate(subjects):
  216. if self.fine_tune_task == 'regression':
  217. target = torch.tensor(
  218. (self.meta_data.loc[self.meta_data['eid'] == subject,
  219. self.target].values[0]
  220. - cont_mean) / cont_std).float()
  221. elif self.fine_tune_task == 'binary_classification':
  222. target = torch.tensor(
  223. self.meta_data.loc[self.meta_data['eid'] == subject,
  224. self.target].values[0])
  225. else: # reconstruction
  226. target = torch.tensor(0)
  227. if self.intermediate_vec == 360:
  228. j = str(subject) + '_20227_2_0'
  229. k = str(subject) + '_20227_3_0'
  230. path_j = os.path.join(self.data_dir, j, 'hcp_mmp1_360_' + j + '.npy')
  231. path_k = os.path.join(self.data_dir, k, 'hcp_mmp1_360_' + k + '.npy')
  232. path_to_fMRIs = path_j if os.path.exists(path_j) else path_k
  233. elif self.intermediate_vec == 400:
  234. path_to_fMRIs = os.path.join(
  235. self.data_dir, str(subject),
  236. 'schaefer_400Parcels_17Networks_' + str(subject) + '.npy')
  237. self.index_l.append((i, subject, path_to_fMRIs, target))
  238. def __len__(self):
  239. return len(self.index_l)
  240. def __getitem__(self, index):
  241. subj, subj_name, path_to_fMRIs, target = self.index_l[index]
  242. # Skip first 20 dummy scans
  243. y = np.load(path_to_fMRIs)[20:20 + self.sequence_length].T
  244. f1, f2 = _find_knee_frequencies(
  245. y, self.TR, self.sequence_length, self.intermediate_vec,
  246. lz_p0=[900, 0.05],
  247. lz_bounds=([0, 0.01], [1200, 0.1]))
  248. ultralow, low, high = _filter_three_bands(
  249. y, self.TR, f1, f2, self.filtering_type)
  250. high = torch.from_numpy(high).T.float()
  251. low = torch.from_numpy(low).T.float()
  252. ultralow = torch.from_numpy(ultralow).T.float()
  253. return {
  254. 'fmri_highfreq_sequence': high,
  255. 'fmri_lowfreq_sequence': low,
  256. 'fmri_ultralowfreq_sequence': ultralow,
  257. 'subject': subj,
  258. 'subject_name': subj_name,
  259. self.target: target,
  260. }
  261. # ─── ABCD ─────────────────────────────────────────────────────────────────────
  262. class ABCD_fMRI_timeseries(BaseDataset):
  263. """Adolescent Brain Cognitive Development (ABCD) dataset.
  264. Targets: sex, ADHD_label, depression, fluid_intelligence, reconstruction."""
  265. TR = 0.8 # seconds
  266. def __init__(self, **kwargs):
  267. self.register_args(**kwargs)
  268. self.data_dir = kwargs.get('abcd_path')
  269. if self.target == 'depression':
  270. self.meta_data = pd.read_csv(
  271. os.path.join(kwargs.get('base_path'), 'data', 'metadata',
  272. 'ABCD_5_1_KSADS_raw_MDD_ANX_CorP_pp_pres_ALL.csv'))
  273. self.meta_data['subjectkey'] = [
  274. i.split('-')[1] for i in self.meta_data['subjectkey']]
  275. self.target = 'MDD_pp'
  276. else:
  277. self.meta_data = pd.read_csv(
  278. os.path.join(kwargs.get('base_path'), 'data', 'metadata',
  279. 'ABCD_matched.csv'))
  280. valid_sub = [i.split('-')[1] for i in os.listdir(self.data_dir)]
  281. if self.target != 'reconstruction':
  282. non_na = self.meta_data[['subjectkey', self.target]].dropna(axis=0)
  283. subjects = list(set(non_na['subjectkey']) & set(valid_sub))
  284. else:
  285. subjects = valid_sub
  286. if self.fine_tune_task == 'regression':
  287. cont_mean = non_na[self.target].mean()
  288. cont_std = non_na[self.target].std()
  289. self.mean = cont_mean
  290. self.std = cont_std
  291. for i, subject in enumerate(subjects):
  292. if self.intermediate_vec == 360:
  293. path_to_fMRIs = os.path.join(
  294. self.data_dir, 'sub-' + subject,
  295. 'hcp_mmp1_sub-' + subject + '.npy')
  296. elif self.intermediate_vec == 400:
  297. path_to_fMRIs = os.path.join(
  298. self.data_dir, 'sub-' + subject,
  299. 'schaefer_sub-' + subject + '.npy')
  300. if self.fine_tune_task == 'regression':
  301. target = torch.tensor(
  302. (self.meta_data.loc[
  303. self.meta_data['subjectkey'] == subject,
  304. self.target].values[0] - cont_mean) / cont_std).float()
  305. elif self.fine_tune_task == 'binary_classification':
  306. target = torch.tensor(
  307. self.meta_data.loc[
  308. self.meta_data['subjectkey'] == subject,
  309. self.target].values[0])
  310. else: # reconstruction
  311. target = torch.tensor(0)
  312. self.index_l.append((i, subject, path_to_fMRIs, target))
  313. def __len__(self):
  314. return len(self.index_l)
  315. def __getitem__(self, index):
  316. subj, subj_name, path_to_fMRIs, target = self.index_l[index]
  317. y = np.load(path_to_fMRIs)[:self.sequence_length].T # [ROI, seq_len]
  318. if self.transfer_learning or self.finetune_test:
  319. pad = 464 - self.sequence_length
  320. f1, f2 = _find_knee_frequencies(
  321. y, self.TR, self.sequence_length, self.intermediate_vec,
  322. lz_p0=[1, 0.05],
  323. lz_bounds=([0, 0.01], [15, 0.1]))
  324. ultralow, low, high = _filter_three_bands(
  325. y, self.TR, f1, f2, self.filtering_type)
  326. if self.transfer_learning or self.finetune_test:
  327. high = F.pad(torch.from_numpy(high),
  328. (pad // 2, pad // 2), 'constant', 0).T.float()
  329. low = F.pad(torch.from_numpy(low),
  330. (pad // 2, pad // 2), 'constant', 0).T.float()
  331. ultralow = F.pad(torch.from_numpy(ultralow),
  332. (pad // 2, pad // 2), 'constant', 0).T.float()
  333. else:
  334. high = torch.from_numpy(high).T.float()
  335. low = torch.from_numpy(low).T.float()
  336. ultralow = torch.from_numpy(ultralow).T.float()
  337. return {
  338. 'fmri_highfreq_sequence': high,
  339. 'fmri_lowfreq_sequence': low,
  340. 'fmri_ultralowfreq_sequence': ultralow,
  341. 'subject': subj,
  342. 'subject_name': subj_name,
  343. self.target: target,
  344. }

datasets.py at commit eed038c, under MIT · at the source

Overview

Authors: Sangyoon Bae1, Junbeom Kwon2, Shinjae Yoo3, Jiook Cha1,4,5
  1. Interdisciplinary Program in Artificial Intelligence, Seoul National University, Seoul, South Korea
  2. Department of Psychology, University of Texas at Austin, Austin, TX USA
  3. Computational Science Initiative, Brookhaven National Laboratory, Shirley, NY USA
  4. Department of Psychology, Seoul National University, Seoul, South Korea
  5. Department of Brain and Cognitive Sciences, Seoul National University, Seoul, South Korea
Institutions: Seoul National University (South Korea); The University of Texas at Austin (United States); Brookhaven National Laboratory (United States)
Journal: Communications biology, volume 9, issue 1, article 963
Dates: received 24 June 2025; accepted 26 March 2026; published online 8 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s42003-026-10011-7 · PMID 42103872 · PMCID PMC13369815 · OpenAlex W7160558314
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), depression (population), autism (population), ADHD (population), cognitive (subfield)
Methods: Spectral & time-frequency, Connectivity, Statistics, Machine learning, Graphs, fMRI & imaging, Physiology & signal measures
Keywords: Learning algorithms, Computational models
MeSH: Attention Deficit Disorder with Hyperactivity*, Autism Spectrum Disorder*, Brain*, Learning*, Adolescent, Brain Mapping, Cognition, Humans, Magnetic Resonance Imaging, Major Depressive Disorder, Nerve Net (* major topic)
Topic: Functional Brain Connectivity Studies (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: Korea Health Industry Development Institute (KHIDI) (HR22C1605); Korea Basic Science Institute (KBSI) (RS-2024-00435727); National Research Foundation of Korea (NRF-2021S1A3A2A02090597, NRF-2021M3E5D2A01022515); U.S. Department of Energy (m4750-2024)
Citations: not cited yet (Europe PMC); 76 references in the paper

Abstract

Understanding how the brain’s nonlinear dynamics give rise to cognition remains a central challenge in neuroscience. Conventional neuroimaging methods assume linearity and stationarity, failing to capture frequency-specific neural computations. We introduce Multi-Band Brain Net (MBBN), a transformer-based framework that models frequency-specific spatiotemporal brain dynamics from fMRI. MBBN integrates biologically grounded frequency decomposition with multi-band self-attention, enabling discovery of frequency-dependent network interactions. Trained on 49,673 individuals across three large-scale cohorts (UK Biobank, Adolescent Brain Cognitive Development Study (ABCD), Autism Brain Imaging Data Exchange (ABIDE)), MBBN achieves state-of-the-art performance in predicting psychiatric and cognitive outcomes—including major depressive disorder (MDD), attention-deficit/hyperactivity disorder (ADHD), and autism spectrum disorder (ASD)—with AUROC improvements of up to 41.36% alongside strong cognitive intelligence prediction. Frequency-resolved analyses reveal disorder-specific signatures: in ADHD, high-frequency fronto-sensorimotor connectivity is attenuated and opercular somatosensory nodes emerge as dynamic hubs; in ASD, orbitofrontal-somatosensory circuits show focal high-frequency disruption alongside enhanced ultra-low-frequency coupling between the temporo-parietal junction and prefrontal cortex. By combining biologically informed frequency decomposition with transformer architecture, MBBN delivers interpretable biomarkers and improved prediction of psychiatric and cognitive traits.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 22 matches between paragraphs and lines of code.

Transconnectome/MBBN

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: eed038cfd162c02461098bae067d0b77de555bec, 9 April 2026
Languages: Python (15), Shell (2), Jupyter (1)
Size: 33 files, 18 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (environment.yaml, requirements.txt), documentation, 1 notebook
Not found: CITATION.cff, tests, continuous integration
Tools: NumPy (15 files), PyTorch (11 files), NiBabel (5 files), pandas (5 files), Nilearn (4 files), scikit-learn (3 files), SciPy (3 files), Matplotlib (2 files), Numba (2 files), scikit-image (2 files), Hugging Face Transformers (2 files), LMFIT (1 file), NetworkX (1 file), Pillow (1 file), PyWavelets (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
20 files

Code availability

The code is available on Github (https://github.com/Transconnectome/MBBN).

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 18 scripts, each with its path and the digest of its content;
  • 22 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

The UKB, ABCD, and ABIDE datasets are publicly available to researchers.

Reproduced under the paper's license (CC BY), from the paper cited above.

Data Availability Statement

The MBBN codebase and environment setup instructions are available in the project repository. Pre-computed data splits are stored in the splits/directory. fMRI data were preprocessed and parcellated according to the HCP-MMP1 (360 parcels) or Schaefer (400 parcels) atlases before model input.

The UKB, ABCD, and ABIDE datasets are publicly available to researchers.

The code is available on Github (https://github.com/Transconnectome/MBBN).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 2 keywords, 11 MeSH terms, 4 funders, 60 references.

Cite

This paper

Bae, S., Kwon, J., Yoo, S., & Cha, J. (2026). Learning brain dynamics across distinct scaling regimes reveals psychiatric signatures. Communications biology, 9(1), 963. https://doi.org/10.1038/s42003-026-10011-7

BibTeX

@article{bae2026learning,
author = {Bae, Sangyoon and Kwon, Junbeom and Yoo, Shinjae and Cha, Jiook},
title = {{Learning brain dynamics across distinct scaling regimes reveals psychiatric signatures}},
journal = {Communications biology},
year = {2026},
month = may,
volume = {9},
number = {1},
pages = {963},
publisher = {Nature Publishing Group},
issn = {2399-3642},
doi = {10.1038/s42003-026-10011-7},
url = {https://doi.org/10.1038/s42003-026-10011-7},
pmid = {42103872},
pmcid = {PMC13369815}
}

RIS

TY - JOUR
AU - Bae, Sangyoon
AU - Kwon, Junbeom
AU - Yoo, Shinjae
AU - Cha, Jiook
TI - Learning brain dynamics across distinct scaling regimes reveals psychiatric signatures
T2 - Communications biology
J2 - Commun Biol
PY - 2026
DA - 2026/05/08
VL - 9
IS - 1
SP - 963
SN - 2399-3642
PB - Nature Publishing Group
DO - 10.1038/s42003-026-10011-7
UR - https://doi.org/10.1038/s42003-026-10011-7
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s42003-026-10011-7",
"type": "article-journal",
"title": "Learning brain dynamics across distinct scaling regimes reveals psychiatric signatures",
"container-title": "Communications biology",
"author": [
{
"family": "Bae",
"given": "Sangyoon"
},
{
"family": "Kwon",
"given": "Junbeom"
},
{
"family": "Yoo",
"given": "Shinjae"
},
{
"family": "Cha",
"given": "Jiook"
}
],
"container-title-short": "Commun Biol",
"volume": "9",
"issue": "1",
"page": "963",
"DOI": "10.1038/s42003-026-10011-7",
"PMID": "42103872",
"PMCID": "PMC13369815",
"ISSN": "2399-3642",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s42003-026-10011-7",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
8
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1002/hbm.70469 [code]
VarCoNet: A Variability-Aware Self-Supervised Framework for Functional Connectome Extraction From Resting-State fMRI.
Journal: Human brain mapping
In common: Hugging Face Transformers, Nilearn, NetworkX, 8 other tools, autism, 2 references
[2] doi:10.3390/brainsci16050455 [code]
Spec-RWKV: A Spectrum-Guided Multi-Scale Recurrent Modeling Framework for Multi-Center Resting-State fMRI-Assisted Diagnosis.
Journal: Brain sciences
In common: PyWavelets, Nilearn, PyTorch, 4 other tools, ADHD, autism, 4 references
[3] doi:10.1162/imag.a.1275 [code]
Glucose metabolism echoes long-range temporal correlations in the human brain.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: NiBabel, scikit-learn, pandas, 3 other tools, 7 references
[4] doi:10.1117/1.nph.13.2.025001 [code]
Surface-based image reconstruction optimization for high-density functional near-infrared spectroscopy.
Journal: Neurophotonics
In common: PyWavelets, Nilearn, NetworkX, 8 other tools
[5] doi:10.1038/s41467-026-76011-7 [code]
Human cortex organizes dynamic co-fluctuations along the sensorimotor-association axis.
Journal: Nature communications
In common: scikit-image, Pillow, NiBabel, 4 other tools, 5 references
[6] doi:10.1162/imag.a.1252 [code]
Does the brain's E:I balance really shape long-range temporal correlations? Lessons learned from 3T MRI.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: PyWavelets, NiBabel, scikit-learn, 4 other tools, 4 references
[7] doi:10.21203/rs.3.rs-9326213/v1 [code]
Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brain
Journal: Research Square (preprint)
In common: Numba, Nilearn, Pillow, 7 other tools, cognitive, 2 references
[8] doi:10.64898/2026.03.09.710558 [code]
Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brain
Journal: bioRxiv (preprint)
In common: Numba, Nilearn, Pillow, 7 other tools, cognitive, 2 references
[9] doi:10.7554/elife.107933 [code]
Modality-agnostic decoding of vision and language from fMRI.
Journal: eLife
In common: Hugging Face Transformers, Nilearn, scikit-image, 8 other tools, cognitive
[10] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Hugging Face Transformers, Numba, NetworkX, 8 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.