OSCR

Brain-CLIPLM: semantic compression for EEG-to-text decoding.

Overview

Authors: Xiaoli Yang1,2, Huiyuan Tian1, Yurui Li1, Jianyu Zhang1, Shijian Li1,2, Gang Pan1,2
  1. College of Computer Science and Technology, Zhejiang University, Hangzhou, China
  2. State Key Lab of Brain-Machine Intelligence, Zhejiang University, Hangzhou, China
Institutions: Zhejiang University (China)
Journal: Frontiers in neuroscience, volume 20, article 1899770
Dates: received 4 June 2026; accepted 6 July 2026; published online 29 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3389/fnins.2026.1899770 · PMID 42591358 · PMCID PMC13461706 · OpenAlex W7171612468
Open access: gold, a free copy (OpenAlex)
Status: data only
Categories: EEG (modality)
Methods: Statistics, Smoothing, state filtering, decompositions, Preprocessing, Machine learning, Physiology & signal measures
Keywords: brain-computer interface, contrastive learning, EEG-to-text decoding, large language models, semantic anchors, semantic compression, sentence reconstruction
Topic: EEG and Brain-Computer Interfaces (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: not cited yet (Europe PMC); 41 references in the paper

Abstract

Decoding natural language from non-invasive electroencephalography (EEG) remains constrained by low signal-to-noise ratio and limited information bandwidth. This raises a central question: can sentence-level language be reliably recovered from such signals? Under realistic information constraints, this direct-recovery assumption may be too strong. We introduce a semantic compression hypothesis: non-invasive EEG may preserve recoverable semantic anchors rather than the full lexical–syntactic form of a sentence. From this perspective, direct sentence reconstruction is overly fine-grained relative to the recoverable information scale of EEG. To address this mismatch, we propose Brain-CLIPLM, a two-stage framework that decomposes EEG-to-text decoding into semantic-anchor recovery and anchor-guided sentence reconstruction. Stage 1 uses contrastive learning to align word-level EEG evidence with a fixed keyword vocabulary and recover ordered semantic anchors. Stage 2 uses a retrieval-grounded large language model with chain-of-thought reasoning prompts to reconstruct sentence meaning from these anchors, following a granularity matching principle that aligns decoding complexity with the recoverable neural information scale. On the combined Zurich Cognitive Language Processing (ZuCo) benchmark, Brain-CLIPLM achieves 67.6% Top-5 and 85.0% Top-25 sentence retrieval accuracy, with the strongest performance at intermediate anchor granularity. Control analyses show that EEG-derived anchors carry sentence-specific information beyond language-model priors. Within the constrained ZuCo sentence pool and fixed keyword-vocabulary settings, these findings suggest that EEG-to-text decoding is better framed as recovering compressed semantic content before anchor-guided sentence reconstruction.

Reproduced under the paper's license (CC BY), from the paper cited above.

Code

The paper links to its data, not to its authors' code: see the Data section.

Tracing map

A tracing map links a paper to the code its authors published: this paper has none, so it has no map.

Data

Datasets cited

Data availability statement

The Zurich Cognitive Language Processing Corpus (ZuCo) benchmark used in this study is constructed from the publicly available ZuCo 1.0 and ZuCo 2.0 datasets. The two datasets are available here: https://osf.io/q3zws/ and https://osf.io/2urht/.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 6 authors, 7 keywords, 27 references.

Cite

This paper

Yang, X., Tian, H., Li, Y., Zhang, J., Li, S., & Pan, G. (2026). Brain-CLIPLM: semantic compression for EEG-to-text decoding. Frontiers in neuroscience, 20, 1899770. https://doi.org/10.3389/fnins.2026.1899770

BibTeX

@article{yang2026brain,
author = {Yang, Xiaoli and Tian, Huiyuan and Li, Yurui and Zhang, Jianyu and Li, Shijian and Pan, Gang},
title = {{Brain-CLIPLM: semantic compression for EEG-to-text decoding}},
journal = {Frontiers in neuroscience},
year = {2026},
month = jul,
volume = {20},
pages = {1899770},
publisher = {Frontiers Media SA},
issn = {1662-4548},
doi = {10.3389/fnins.2026.1899770},
url = {https://doi.org/10.3389/fnins.2026.1899770},
pmid = {42591358},
pmcid = {PMC13461706}
}

RIS

TY - JOUR
AU - Yang, Xiaoli
AU - Tian, Huiyuan
AU - Li, Yurui
AU - Zhang, Jianyu
AU - Li, Shijian
AU - Pan, Gang
TI - Brain-CLIPLM: semantic compression for EEG-to-text decoding
T2 - Frontiers in neuroscience
J2 - Front Neurosci
PY - 2026
DA - 2026/07/29
VL - 20
SP - 1899770
SN - 1662-4548
PB - Frontiers Media SA
DO - 10.3389/fnins.2026.1899770
UR - https://doi.org/10.3389/fnins.2026.1899770
LA - en
ER -

CSL-JSON

{
"id": "10.3389/fnins.2026.1899770",
"type": "article-journal",
"title": "Brain-CLIPLM: semantic compression for EEG-to-text decoding",
"container-title": "Frontiers in neuroscience",
"author": [
{
"family": "Yang",
"given": "Xiaoli"
},
{
"family": "Tian",
"given": "Huiyuan"
},
{
"family": "Li",
"given": "Yurui"
},
{
"family": "Zhang",
"given": "Jianyu"
},
{
"family": "Li",
"given": "Shijian"
},
{
"family": "Pan",
"given": "Gang"
}
],
"container-title-short": "Front Neurosci",
"volume": "20",
"page": "1899770",
"DOI": "10.3389/fnins.2026.1899770",
"PMID": "42591358",
"PMCID": "PMC13461706",
"ISSN": "1662-4548",
"publisher": "Frontiers Media SA",
"URL": "https://doi.org/10.3389/fnins.2026.1899770",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
29
]
]
}
}

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10265-1 [code]
The shape of attention reflects flexible filtering of natural speech modulations.
Journal: Communications biology
In common: EEG, 3 references
[2] doi:10.1162/nol.a.264 [code]
The Temporal Dynamics of the Labeling Algorithm During Natural Language Comprehension: Neural Evidence for Phrase Grammatical Type Generation.
Journal: Neurobiology of language (Cambridge, Mass.)
In common: EEG, 3 references
[3] doi:10.1371/journal.pbio.3003924
Endogenous auditory and motor brain rhythms predict individual speech tracking.
Journal: PLoS biology
In common: 3 references
[4] doi:10.3390/s26103212
Imagined Speech Brain-Computer Interface: A Task-Oriented Review of Neural Decoding.
Journal: Sensors (Basel, Switzerland)
In common: EEG, 2 references
[5] doi:10.1038/s41598-025-29587-x [code]
Evaluating EEG-to-text models through noise-based performance analysis
Journal: n/a
In common: EEG, 2 references
[6] doi:10.3758/s13415-026-01442-0 [code]
Empathy for pain persists across live two-way video interactions and viewing of prerecorded videos.
Journal: Cognitive, affective & behavioral neuroscience
In common: EEG, 2 references
[7] doi:10.1038/s42003-026-10394-7 [code]
Cognitive load weakens neural speech tracking without altering response timing.
Journal: Communications biology
In common: EEG, 2 references
[8] doi:10.1111/cogs.70220 [code]
Shared Neural Computations for Syntactic and Morphological Structures: Evidence From Mandarin Chinese.
Journal: Cognitive science
In common: EEG, 2 references
[9] doi:10.1002/aur.70312 [code]
Aberrant Neural Entrainment to Word-Level Speech Patterns in Fragile X Syndrome: Evidence for a Statistical Learning Deficit.
Journal: Autism research : official journal of the International Society for Autism Research
In common: EEG, 2 references
[10] doi:10.1162/nol.a.271 [code]
Compositional Complexity in Text and Images.
Journal: Neurobiology of language (Cambridge, Mass.)
In common: 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.