OSCR

Compositional Complexity in Text and Images.

Code ↔ Paper

12 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 12 matches
  1. [1] § MATERIALS AND METHODS › Experimental Procedure ↔ src/compositionality_study/experiment/mri/convert_hf_ds_to_local_files.py, lines 214–270 · score 0.82 · inter stimulus interval, blank trials, randomized, ISI, duration, split
  2. [2] § MATERIALS AND METHODS › Experimental Procedure ↔ src/compositionality_study/experiment/mri/generate_psychopy_files.py, lines 206–354 · score 0.80 · inter stimulus interval, frame rate, practice stimuli, PsychoPy, trigger, duration
  3. [3] § MATERIALS AND METHODS › Stimuli › Stimulus selection ↔ src/compositionality_study/data/select_stimuli.py, lines 134–208 · score 0.75 · propensity score matching, PsmPy, duplications, subset, filtered, graph
  4. [4] § MATERIALS AND METHODS › Image Preprocessing ↔ src/compositionality_study/models/univariate_analysis.py, lines 61–167 · score 0.72 · motion regressors, drift model, cosine, GLMs, smoothing, scans
  5. [5] § MATERIALS AND METHODS › Statistical Analyses ↔ src/compositionality_study/models/mvpa_decoder.py, lines 194–285 · score 0.68 · modality decoding, radius, stratified, fold, searchlight, linear
  6. [6] § MATERIALS AND METHODS › Stimuli › Textual compositional complexity ↔ src/compositionality_study/utils.py, lines 279–325 · score 0.61 · dependency parsing, graph depth, AMR graph, trees, edges, nodes
  7. [7] § MATERIALS AND METHODS › Stimuli › Textual compositional complexity ↔ src/compositionality_study/visualization/visualize_selected_stimuli.py, lines 227–307 · score 0.61 · dependency parsing, graph depth, AMR graph, attributes, trees, model
  8. [8] § MATERIALS AND METHODS › Image Preprocessing ↔ src/compositionality_study/models/post_hoc.py, lines 58–123 · score 0.57 · drift model, cosine, motion, GLMs, smoothing, confounds
  9. [9] § MATERIALS AND METHODS › Stimuli › Textual compositional complexity ↔ src/compositionality_study/utils.py, lines 279–325 · score 0.55 · dependency parsing, AMR graph, rooted, trees, edges, sentence
  10. [10] § MATERIALS AND METHODS › Stimuli › Visual compositional complexity ↔ src/compositionality_study/utils.py, lines 227–257 · score 0.53 · graph depth, AMR graphs, longest, shortest, edges, nodes
  11. [11] § RESULTS › Multivariate Pattern Analysis of Compositional Complexity ↔ src/compositionality_study/models/estimate_betas.py, lines 82–128 · score 0.51 · extra confounds, aspect ratio, beta, nodes, model, AMR
  12. [12] § RESULTS › Univariate Analysis of Compositional Complexity ↔ src/compositionality_study/models/univariate_analysis.py, lines 536–594 · score 0.50 · functional localization mask, conjunction, row, confounds, frame, Univariate

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 510 lines · 18 KB · MIT · 3 matches

  1. """Utils for the compositionality project."""
  2. from typing import Dict, Iterable, List, Optional, Tuple, Union
  3. from pathlib import Path
  4. import amrlib
  5. import networkx as nx
  6. import nibabel as nib
  7. import numpy as np
  8. import pandas as pd
  9. import penman
  10. from datasets import load_dataset
  11. from loguru import logger
  12. from nilearn import image, plotting
  13. from nilearn.glm.thresholding import threshold_stats_img
  14. from penman.exceptions import DecodeError
  15. from PIL import Image, ImageOps
  16. from spacy.language import Language # type: ignore
  17. from spacy.tokens import Doc, Token
  18. from compositionality_study.constants import HF_DATASET_NAME
  19. # Set up the spacy amrlib extension
  20. amrlib.setup_spacy_extension()
  21. _COCO_DF = None
  22. DEFAULT_VOXEL_P = 0.001
  23. DEFAULT_CLUSTER_THRESHOLD = 50
  24. def get_coco_df() -> pd.DataFrame:
  25. """Load and cache the COCO dataset."""
  26. global _COCO_DF
  27. if _COCO_DF is None:
  28. ds = load_dataset(HF_DATASET_NAME, split="train")
  29. df = ds.to_pandas()
  30. if isinstance(df, pd.DataFrame):
  31. _COCO_DF = df
  32. else:
  33. raise ValueError("Expected DataFrame")
  34. return _COCO_DF # type: ignore
  35. def get_stimulus_features_lookup(coco_df: pd.DataFrame) -> Tuple[pd.DataFrame, pd.DataFrame]:
  36. """Create optimized lookup tables for text and image features."""
  37. txt_df = (
  38. coco_df.drop_duplicates("sentences_raw").set_index("sentences_raw")
  39. if "sentences_raw" in coco_df.columns
  40. else pd.DataFrame()
  41. )
  42. img_df = (
  43. coco_df.drop_duplicates("cocoid").set_index("cocoid")
  44. if "cocoid" in coco_df.columns
  45. else pd.DataFrame()
  46. )
  47. return txt_df, img_df
  48. def get_stimulus_data(
  49. modality: Optional[str],
  50. stimulus: Optional[str],
  51. cocoid: Union[str, float, int, None],
  52. txt_df: pd.DataFrame,
  53. img_df: pd.DataFrame,
  54. ) -> Union[pd.Series, None]: # type: ignore
  55. """Retrieve features for a single stimulus event."""
  56. if modality == "text" and stimulus in txt_df.index:
  57. res = txt_df.loc[stimulus]
  58. return res if isinstance(res, pd.Series) else res.iloc[0]
  59. elif modality == "image":
  60. if pd.notna(cocoid):
  61. if cocoid in img_df.index:
  62. res = img_df.loc[cocoid]
  63. return res if isinstance(res, pd.Series) else res.iloc[0]
  64. try:
  65. # Handle potential float/string mismatches
  66. int_cid = int(cocoid) # type: ignore
  67. if int_cid in img_df.index:
  68. res = img_df.loc[int_cid]
  69. return res if isinstance(res, pd.Series) else res.iloc[0]
  70. except (ValueError, TypeError):
  71. pass
  72. return None
  73. def derive_run_id(bold_file: Path) -> str:
  74. """Strip space/desc tokens to get the BIDS run identifier."""
  75. stem = bold_file.name.replace(".nii.gz", "").replace(".nii", "")
  76. parts: List[str] = []
  77. for token in stem.split("_"):
  78. if token.startswith("space-") or token.startswith("desc-"):
  79. break
  80. parts.append(token)
  81. return "_".join(parts)
  82. def load_events(events_tsv: Path) -> pd.DataFrame:
  83. """Load events and normalize the condition column name."""
  84. events = pd.read_csv(events_tsv, sep="\t")
  85. modality = events["modality"].astype(str).str.strip().str.lower()
  86. complexity = events["complexity"].astype(str).str.strip().str.lower()
  87. events["trial_type"] = modality + "_" + complexity
  88. if "trial_type" in events.columns:
  89. events = events.dropna(subset=["trial_type"])
  90. events["trial_type"] = events["trial_type"].astype(str)
  91. return events
  92. def find_gm_mask(fmriprep_dir: Path, subject: str) -> Path:
  93. """Pick the GM probseg (assuming MNI152NLin2009cAsym space)."""
  94. anat_dir = fmriprep_dir / f"sub-{subject}" / "ses-01" / "anat"
  95. candidates = sorted(anat_dir.glob("*space-MNI152NLin2009cAsym*_label-GM_probseg.nii.gz"))
  96. return candidates[0]
  97. def collect_runs(
  98. fmriprep_dir: Path, bids_dir: Path, subject: str, max_runs: int | None = None
  99. ) -> List[Tuple[Path, Path, Path]]:
  100. """Pair each preprocessed BOLD run with matching events and confounds."""
  101. bold_files = sorted((fmriprep_dir / f"sub-{subject}").glob("**/*_desc-preproc_bold.nii.gz"))
  102. runs: List[Tuple[Path, Path, Path]] = []
  103. if max_runs is not None:
  104. logger.info(f"Debugging: limiting to {max_runs} runs for subject {subject}")
  105. bold_files = bold_files[:max_runs]
  106. for bold_file in bold_files:
  107. run_id = derive_run_id(bold_file)
  108. rel_dir = bold_file.parent.relative_to(fmriprep_dir / f"sub-{subject}")
  109. events_file = bids_dir / f"sub-{subject}" / rel_dir / f"{run_id}_events.tsv"
  110. confounds_file = bold_file.parent / f"{run_id}_desc-confounds_timeseries.tsv"
  111. runs.append((bold_file, events_file, confounds_file))
  112. return runs
  113. def load_stimulus_confounds(events: pd.DataFrame, n_scans: int, tr: float) -> pd.DataFrame:
  114. """Generate extra confounds based on stimulus properties."""
  115. coco_df = get_coco_df()
  116. text_cols = ["sentence_length", "amr_n_nodes"]
  117. img_cols = ["coco_a_nodes", "ic_score", "aspect_ratio", "coco_person"]
  118. all_cols = text_cols + img_cols
  119. confounds = pd.DataFrame(0.0, index=range(n_scans), columns=all_cols)
  120. txt_df, img_df = get_stimulus_features_lookup(coco_df)
  121. for _, row in events.iterrows():
  122. data = get_stimulus_data(
  123. modality=row.get("modality"),
  124. stimulus=row.get("stimulus"),
  125. cocoid=row.get("cocoid"),
  126. txt_df=txt_df,
  127. img_df=img_df,
  128. )
  129. if data is None:
  130. continue
  131. start_tr = int(round(row["onset"] / tr))
  132. n_trs = int(round(row["duration"] / tr))
  133. end_tr = min(start_tr + n_trs, n_scans)
  134. if start_tr >= n_scans:
  135. continue
  136. cols = text_cols if row.get("modality") == "text" else img_cols
  137. valid_cols = [c for c in cols if c in data]
  138. if valid_cols:
  139. confounds.iloc[
  140. start_tr:end_tr, confounds.columns.get_indexer(valid_cols) # type: ignore
  141. ] = data[valid_cols].values
  142. return confounds
  143. def conjunction_map(map_a: Path, map_b: Path) -> Path:
  144. """Compute a simple conjunction (logical AND) of two thresholded maps."""
  145. img_a = image.math_img("img != 0", img=map_a)
  146. img_b = image.math_img("img != 0", img=map_b)
  147. conj = image.math_img("img1 * img2", img1=img_a, img2=img_b)
  148. out_file = map_a.with_name("conjunction_img_txt.nii.gz")
  149. save_brain_map(conj, out_file)
  150. return out_file
  151. def split_runs_by_session(runs: Iterable[Tuple[Path, Path, Path]]) -> Dict[str, List[Tuple[Path, Path, Path]]]:
  152. """Group runs by session token for downstream localizer logic."""
  153. grouped: Dict[str, List[Tuple[Path, Path, Path]]] = {}
  154. for bold, events, confounds in runs:
  155. run_id = derive_run_id(bold)
  156. session = run_id.split("_")[1]
  157. grouped.setdefault(session, []).append((bold, events, confounds))
  158. return grouped
  159. def get_aspect_ratio(filepath: str):
  160. """Get the aspect ratio of the image from the local directory.
  161. :param filepath: The path to the image file
  162. :type filepath: str
  163. :return: The aspect ratio of the image
  164. :rtype: float
  165. """
  166. # Load the image
  167. try:
  168. with Image.open(filepath) as img:
  169. width, height = img.size
  170. return width / height
  171. except Exception as e: # noqa
  172. return 0.0
  173. def get_amr_graph_depth(
  174. amr_graph: str,
  175. return_graph=False,
  176. ) -> Union[int, Tuple[int, nx.DiGraph]]:
  177. """Get the depth of the AMR graph for a given example.
  178. :param amr_graph: The AMR graph to get the depth for (output of a spacy doc._.to_amr()[0] call)
  179. :type amr_graph: str
  180. :param return_graph: Whether to return the networkx graph, defaults to False
  181. :type return_graph: bool, optional
  182. :return: The maximum "depth" of the AMR graph (longest shortest path)
  183. :rtype: int
  184. """
  185. # Convert to a Penman graph (with de-inverted edges)
  186. penman_graph = penman.decode(amr_graph)
  187. # Convert to a nx graph, first initialize the nx graph
  188. nx_graph = nx.DiGraph()
  189. # Add edges
  190. for e in penman_graph.edges():
  191. nx_graph.add_edge(e.source, e.target)
  192. # Get the characteristic path length of the graph
  193. amr_graph_depth = (
  194. max([max(nx.shortest_path_length(nx_graph, source=n).values()) for n in nx_graph.nodes()])
  195. if nx.number_of_nodes(nx_graph) > 0
  196. else 0
  197. )
  198. if return_graph:
  199. return amr_graph_depth, nx_graph
  200. else:
  201. return amr_graph_depth
  202. def walk_tree(
  203. node: Token,
  204. depth: int,
  205. ) -> int:
  206. """Walk the dependency parse tree and return the maximum depth.
  207. :param node: The current node in the tree
  208. :type node: spacy.tokens.Token
  209. :param depth: The current depth in the tree
  210. :type depth: int
  211. :return: The maximum depth in the tree
  212. :rtype: int
  213. """
  214. if node.n_lefts + node.n_rights > 0:
  215. return max(walk_tree(child, depth + 1) for child in node.children)
  216. else:
  217. return depth
  218. def derive_text_depth_features(
  219. examples: Dict[str, List],
  220. nlp: Language,
  221. ) -> Dict[str, List]:
  222. """Get the depth of the dep parse tree, number of verbs and "depth" of the AMR graph of an example caption.
  223. The AMR model needs to be downloaded separately, see https://github.com/bjascob/amrlib-models.
  224. :param examples: A batch of hf dataset examples
  225. :type examples: Dict[str, List]
  226. :param nlp: Spacy pipeline to use, initialized using nlp = spacy.load("en_core_web_trf")
  227. :type nlp: spacy.language.Language
  228. :return: The batch with the added features
  229. :rtype: Dict[str, List]
  230. """
  231. result: Dict = {
  232. "parse_tree_depth": [],
  233. "n_verbs": [],
  234. "amr_graph_depth": [],
  235. "amr_graph": [],
  236. "amr_n_nodes": [],
  237. "amr_n_edges": [],
  238. }
  239. doc_batched = nlp.pipe(examples["sentences_raw"])
  240. for doc in doc_batched:
  241. # Also derive the AMR graph for the caption and derive its depth
  242. amr_graph = doc._.to_amr()[0] # type: ignore
  243. try:
  244. amr_depth, amr_graph_obj = get_amr_graph_depth(amr_graph, return_graph=True) # type: ignore
  245. n_nodes = nx.number_of_nodes(amr_graph_obj)
  246. n_edges = nx.number_of_edges(amr_graph_obj)
  247. amr_graph_arr = nx.to_numpy_array(amr_graph_obj)
  248. except DecodeError:
  249. amr_depth = 0
  250. n_nodes = 0
  251. n_edges = 0
  252. amr_graph_arr = nx.to_numpy_array(nx.DiGraph())
  253. # Determine the depth of the dependency parse tree
  254. result["parse_tree_depth"].append(walk_tree(next(doc.sents).root, 0))
  255. result["n_verbs"].append(len([token for token in doc if token.pos_ == "VERB"]))
  256. result["amr_graph_depth"].append(amr_depth)
  257. result["amr_graph"].append(amr_graph_arr)
  258. result["amr_n_nodes"].append(n_nodes)
  259. result["amr_n_edges"].append(n_edges)
  260. return examples | result
  261. def dependency_parse_to_nx(
  262. sents: List[Doc],
  263. ):
  264. """Convert spaCy sentence objects into a NetworkX directed graph representing the dependency parse tree.
  265. :param sents: A list of spaCy sentence objects
  266. :type sents: List[spacy.tokens.Doc]
  267. :return: A NetworkX directed graph representing the dependency parse tree
  268. :rtype: nx.DiGraph
  269. """
  270. # Create a directed graph
  271. graph = nx.DiGraph()
  272. # Iterate over sentences
  273. for sent in sents:
  274. for token in sent:
  275. # Add node for the token with attributes
  276. graph.add_node(token.i, text=token.text, pos=token.pos_, tag=token.tag_)
  277. # Add edge from head to child (if not the root token)
  278. if token.head != token: # Avoid self-loop for the root
  279. graph.add_edge(token.head.i, token.i, dep=token.dep_)
  280. return graph
  281. def flatten_examples(
  282. examples: Dict[str, List],
  283. flatten_col_names: List[str] = ["sentences_raw", "sentids"],
  284. ) -> Dict[str, List]:
  285. """Flattens the examples in the dataset.
  286. :param examples: The examples to flatten
  287. :type examples: Dict[str, List]
  288. :param flatten_col_names: The column names to flatten, defaults to ["sentences_raw", "sentids"]
  289. :type flatten_col_names: List[str]
  290. :return: The flattened examples
  291. :rtype: Dict[str, List]
  292. """
  293. flattened_data = {}
  294. number_of_sentences = [len(sents) for sents in examples[flatten_col_names[0]]]
  295. for key, value in examples.items():
  296. if key in flatten_col_names:
  297. flattened_data[key] = [sent for sent_list in value for sent in sent_list]
  298. else:
  299. flattened_data[key] = np.repeat(value, number_of_sentences).tolist()
  300. return flattened_data
  301. def apply_gamma_correction(
  302. image: Image.Image,
  303. target_mean=128.0,
  304. ) -> Image.Image:
  305. """Apply gamma correction to an image.
  306. :param image: The image to apply gamma correction to
  307. :type image: PIL.Image.Image
  308. :param target_mean: The target mean brightness of the image, defaults to 128.0
  309. :type target_mean: float, optional
  310. :return: The image with gamma correction applied
  311. :rtype: PIL.Image.Image
  312. """
  313. # Convert the PIL image to a numpy array
  314. img_array = np.array(image)
  315. # Calculate the current mean brightness of the image
  316. current_mean = np.mean(img_array)
  317. # Calculate the gamma value to adjust the mean to the target mean
  318. # Avoid division by zero
  319. if current_mean > 0:
  320. gamma = np.log(target_mean) / np.log(current_mean)
  321. # Apply gamma correction to the image
  322. corrected_image = ImageOps.autocontrast(image, cutoff=gamma) # type: ignore
  323. return corrected_image
  324. else:
  325. return image
  326. def save_brain_map(
  327. img: Union[nib.nifti1.Nifti1Image, object],
  328. output_path: Union[str, Path],
  329. is_z_map: bool = False,
  330. voxel_p: float = DEFAULT_VOXEL_P,
  331. cluster_threshold: int = DEFAULT_CLUSTER_THRESHOLD,
  332. make_surface_plot: bool = True,
  333. surface_plot_kwargs: Optional[Dict] = None,
  334. sided: str = "positive",
  335. ) -> Optional[nib.nifti1.Nifti1Image]:
  336. """Save a brain map to disk, generate mosaic plots, and cluster tables.
  337. For z-score/stat maps (``is_z_map=True``), apply one-sided FPR thresholding
  338. with ``threshold_stats_img`` (using ``voxel_p`` and ``cluster_threshold``).
  339. Set ``sided="positive"`` to keep only positive effects, ``sided="negative"``
  340. for negative effects, or ``sided="both"`` to save both directions (files
  341. suffixed with ``_pos_thr`` and ``_neg_thr``).
  342. :param img: The brain map image
  343. :param output_path: Where to save the image
  344. :param is_z_map: Whether the image is a z-/stat-map that should be thresholded
  345. :param voxel_p: Two-sided voxel-wise p-value used for thresholding
  346. :param make_surface_plot: Whether to save a surface plot
  347. :param surface_plot_kwargs: Extra args for ``generate_surface_plots``
  348. :param sided: Direction for one-sided thresholding (positive, negative, both)
  349. :returns: Thresholded positive image if available, otherwise None
  350. """
  351. output_path = Path(output_path)
  352. output_path.parent.mkdir(parents=True, exist_ok=True)
  353. def _base_path(path: Path) -> str:
  354. str_path = str(path)
  355. if str_path.endswith(".nii.gz"):
  356. return str_path[:-7]
  357. if str_path.endswith(".nii"):
  358. return str_path[:-4]
  359. return str(path.with_suffix(""))
  360. def _plot(target_img: object, target_path: Path, cmap: Optional[str] = None) -> None:
  361. base = _base_path(target_path)
  362. try:
  363. display = plotting.plot_stat_map(
  364. target_img,
  365. display_mode="mosaic",
  366. cut_coords=5,
  367. cmap=cmap, # type: ignore
  368. )
  369. display.savefig(f"{base}_mosaic.png") # type: ignore
  370. display.close() # type: ignore
  371. except Exception as e:
  372. logger.error(f"Failed to plot brain map: {e}")
  373. # Save the Nifti image
  374. if hasattr(img, "to_filename"):
  375. img.to_filename(output_path) # type: ignore
  376. else:
  377. nib.save(img, output_path) # type: ignore
  378. _plot(img, output_path)
  379. thresholded_img = None
  380. if is_z_map:
  381. try:
  382. do_pos = sided in ("positive", "both")
  383. do_neg = sided in ("negative", "both")
  384. def _save_thr(target_img: object, suffix: str, flip_back: bool = False, cmap: Optional[str] = None) -> object:
  385. thr_img, _ = threshold_stats_img(
  386. target_img,
  387. alpha=voxel_p,
  388. height_control="fpr",
  389. cluster_threshold=cluster_threshold,
  390. two_sided=False,
  391. )
  392. if flip_back:
  393. thr_img = image.math_img("-img", img=thr_img)
  394. if str(output_path).endswith(".nii.gz"):
  395. thr_path = output_path.with_name(output_path.name.replace(".nii.gz", f"_{suffix}.nii.gz"))
  396. else:
  397. thr_path = output_path.with_name(f"{output_path.stem}_{suffix}{output_path.suffix}")
  398. thr_img.to_filename(thr_path) # type: ignore
  399. _plot(thr_img, thr_path, cmap=cmap)
  400. return thr_img
  401. if do_pos:
  402. thresholded_img = _save_thr(img, "pos_thr", cmap="Reds")
  403. if do_neg:
  404. _save_thr(image.math_img("-img", img=img), "neg_thr", flip_back=True, cmap="Blues_r")
  405. except Exception as e:
  406. logger.error(f"Failed to threshold brain map: {e}")
  407. # Save surface plot
  408. if make_surface_plot:
  409. from compositionality_study.visualization.visualize_brain_maps import generate_surface_plots
  410. surface_plot_kwargs = surface_plot_kwargs or {}
  411. try:
  412. generate_surface_plots(
  413. nii_file=output_path,
  414. output_dir=output_path.parent,
  415. **surface_plot_kwargs
  416. )
  417. except Exception as e:
  418. logger.error(f"Failed to generate surface plot: {e}")
  419. return thresholded_img # type: ignore

utils.py at commit bb03182, under MIT · at the source

Overview

  1. Laboratory for Cognitive Neurology, Department of Neurosciences, Leuven Brain Institute, KU Leuven, Leuven, Belgium
  2. Language Intelligence and Information Retrieval Lab, Department of Computer Science, KU Leuven, Leuven, Belgium
Institutions: KU Leuven (Belgium)
Journal: Neurobiology of language (Cambridge, Mass.), volume 7, article NOL.a.271
Dates: received 27 February 2025; accepted 28 April 2026; published online 9 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1162/nol.a.271 · PMID 42644209 · PMCID PMC13506228 · OpenAlex W7160494749
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: fMRI (modality)
Methods: Statistics, Machine learning, Connectivity, fMRI & imaging, Single-unit activity, calcium imaging, Smoothing, state filtering, decompositions
Keywords: compositionality, functional magnetic resonance imaging, multivariate pattern analysis
Topic: Neurobiology of Language and Bilingualism (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: Fonds Wetenschappelijk Onderzoek (10.13039/501100003130, 1154623N, 501100003130, 1247821N); KU Leuven (501100004040, C14/21/109)
Citations: not cited yet (Europe PMC); 78 references in the paper

Abstract

Compositionality enables us to derive the meaning of a complex whole from the syntax and semantics of its individual parts. While compositional processing has primarily been studied in language, similar principles apply to visual scenes, in which meaning can be derived from their individual parts and the relationships between them. In this study, we therefore aimed to explore whether there is a shared neural basis for compositional processing in text and images. Based on abstract meaning representation graphs for text, commonly used to capture “who is doing what to whom”, and action graph annotations for visual scenes, we defined an analogous graph depth-based notion of compositional complexity. By conducting a functional magnetic resonance imaging experiment in which participants view text-image pairs from the Common Objects in Context-Actions data set, we aimed to identify brain activity patterns related to compositional processing through a combination of univariate and multivariate pattern analysis. Specifically, we hypothesized that there are shared brain regions—such as the pars opercularis, pars triangularis or anterior temporal lobe—involved in processing compositional complexity across text and images. Alternatively, there might exist distinct regions for processing compositional complexity in text and images, without any substantial cross-modal overlap. Our results provided no evidence in support of either hypothesis: no significant neural responses to compositional complexity were observed, in either a shared or modality-specific manner. Given the rigorous control of confounds in our study design, the operationalization of compositional complexity itself may underlie the null result.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 12 matches between paragraphs and lines of code.

lcn-kul/compositionality-study

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: bb03182371decdfede04bfac39e5dc335e2a789d, 22 April 2026
Languages: Python (31)
Size: 52 files, 31 scripts
Software Heritage: archived
Found in: “DATA AND CODE AVAILABILITY STATEMENTS”
Holds: README, license file, CITATION.cff, environment (poetry.lock, pyproject.toml, setup.cfg, setup.py, uv.lock), continuous integration, documentation
Not found: tests
Tools: pandas (15 files), NumPy (12 files), Pillow (6 files), Matplotlib (5 files), Nilearn (5 files), NetworkX (4 files), NiBabel (4 files), SciPy (4 files), seaborn (3 files), GLMsingle (2 files), OpenCV (2 files), PyTorch (2 files), Pingouin (1 file), PsychoPy (1 file), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
33 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 31 scripts, each with its path and the digest of its content;
  • 12 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data and code availability statements

The approved Stage 1 protocol, laboratory log, stimuli and behavioral results can be found on the Open Science Framework: https://osf.io/ta45d/overview?view_only=439003edb48c4053ab464d000894fa3e. Furthermore, the pre-processed study data supporting the reported analysis is available at Zenodo: https://zenodo.org/records/18831683 and https://zenodo.org/records/18832298. All code used for the pre-registered analyses is publicly available as well: https://github.com/lcn-kul/compositionality-study.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Funding: added Fonds Wetenschappelijk Onderzoek: 10.13039/501100003130, 1154623N, 501100003130, 1247821N; KU Leuven: 501100004040, C14/21/109

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 6 authors, 3 keywords, 71 references.

Cite

This paper

Balabin, H., Liuzzi, A. G., Statz, K., Moens, M.-F., Dupont, P., & Vandenberghe, R. (2026). Compositional Complexity in Text and Images. Neurobiology of language (Cambridge, Mass.), 7, NOL.a.271. https://doi.org/10.1162/nol.a.271

BibTeX

@article{balabin2026compositional,
author = {Balabin, Helena and Liuzzi, Antonietta Gabriella and Statz, Kevin and Moens, Marie-Francine and Dupont, Patrick and Vandenberghe, Rik},
title = {{Compositional Complexity in Text and Images}},
journal = {Neurobiology of language (Cambridge, Mass.)},
year = {2026},
month = jul,
volume = {7},
pages = {NOL.a.271},
publisher = {MIT Press},
issn = {2641-4368},
doi = {10.1162/nol.a.271},
url = {https://doi.org/10.1162/nol.a.271},
pmid = {42644209},
pmcid = {PMC13506228}
}

RIS

TY - JOUR
AU - Balabin, Helena
AU - Liuzzi, Antonietta Gabriella
AU - Statz, Kevin
AU - Moens, Marie-Francine
AU - Dupont, Patrick
AU - Vandenberghe, Rik
TI - Compositional Complexity in Text and Images
T2 - Neurobiology of language (Cambridge, Mass.)
J2 - Neurobiol Lang (Camb)
PY - 2026
DA - 2026/07/09
VL - 7
SP - NOL.a.271
SN - 2641-4368
PB - MIT Press
DO - 10.1162/nol.a.271
UR - https://doi.org/10.1162/nol.a.271
LA - en
ER -

CSL-JSON

{
"id": "10.1162/nol.a.271",
"type": "article-journal",
"title": "Compositional Complexity in Text and Images",
"container-title": "Neurobiology of language (Cambridge, Mass.)",
"author": [
{
"family": "Balabin",
"given": "Helena"
},
{
"family": "Liuzzi",
"given": "Antonietta Gabriella"
},
{
"family": "Statz",
"given": "Kevin"
},
{
"family": "Moens",
"given": "Marie-Francine"
},
{
"family": "Dupont",
"given": "Patrick"
},
{
"family": "Vandenberghe",
"given": "Rik"
}
],
"container-title-short": "Neurobiol Lang (Camb)",
"volume": "7",
"page": "NOL.a.271",
"DOI": "10.1162/nol.a.271",
"PMID": "42644209",
"PMCID": "PMC13506228",
"ISSN": "2641-4368",
"publisher": "MIT Press",
"URL": "https://doi.org/10.1162/nol.a.271",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
9
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.7554/elife.107933 [code]
Modality-agnostic decoding of vision and language from fMRI.
Journal: eLife
In common: Nilearn, OpenCV, Pillow, 8 other tools, fMRI, 7 references
[2] doi:10.1038/s41597-026-07248-6 [code]
A large-scale fMRI dataset for vision-language semantic association.
Journal: Scientific data
In common: GLMsingle, Nilearn, OpenCV, 8 other tools, fMRI, 4 references
[3] doi:10.1523/jneurosci.0038-26.2026 [code]
Multidimensional Feature Tuning in Category Selective Areas of Human Visual Cortex.
Journal: The Journal of neuroscience : the official journal of the Society for Neuroscience
In common: Pillow, NiBabel, PyTorch, 6 other tools, fMRI, 6 references
[4] doi:10.1162/imag.a.1309 [code]
Probing the content of semantic representations in body-selective regions.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: OpenCV, Pillow, NiBabel, 7 other tools, 5 references
[5] doi:10.1038/s41467-026-76098-y [code]
A single computational objective can produce specialization of streams in visual cortex.
Journal: Nature communications
In common: Pillow, NiBabel, PyTorch, 6 other tools, 5 references
[6] doi:10.1038/s42003-026-10169-0 [code]
Shared representations in brains and models reveal a two-route cortical organization during scene perception.
Journal: Communications biology
In common: Pillow, NiBabel, PyTorch, 6 other tools, 5 references
[7] doi:10.1162/imag.a.1286 [code]
Behavioral imitation with artificial neural networks leads to personalized models of brain dynamics during videogame play.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Nilearn, Pillow, NiBabel, 7 other tools, fMRI, 4 references
[8] doi:10.1162/imag.a.1256 [code]
Gamer in the scanner: Event-related analysis of fMRI activity during retro videogame play guided by automated annotations of game content.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Nilearn, Pillow, NiBabel, 7 other tools, fMRI, 4 references
[9] doi:10.1167/jov.26.5.7 [code]
Representations in vision and language converge in a shared, multidimensional space of perceived similarities.
Journal: Journal of vision
In common: Nilearn, OpenCV, Pillow, 8 other tools, 2 references
[10] doi:10.1007/s12021-026-09803-3 [code]
NeuroFusion: A Unified Framework for Generalized Visual Stimulus Decoding from fMRI Across Datasets and Subjects.
Journal: Neuroinformatics
In common: Nilearn, Pillow, NiBabel, 6 other tools, fMRI, 3 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.