OSCR

From thought to language: Comparing schizophrenia spectrum disorders and Wernicke's aphasia with machine learning and LLMs.

Code ↔ Paper

9 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 9 matches
  1. [1] § Methods › Feature extraction › Syntactic features ↔ udstyle.py, lines 1–33 · score 0.97 · Lexical density, Normalized dependency distance, clause length, nominal modifiers, adjacent dependencies, complexity metrics
  2. [2] § Methods › Feature extraction › Syntactic features ↔ udstyle.py, lines 1–33 · score 0.62 · tag frequencies, complexity metrics, UDstyle, POS, lexical
  3. [3] § Results › Supervised machine learning using extracted features › Feature importance ↔ udstyle.py, lines 256–345 · score 0.59 · possessive nominal modifier, conjunction, pronoun, nmod, nouns, sentence
  4. [4] § Methods › Models for classification › Zero-shot classification. ↔ zeroshot/zeroshot.py, lines 1–48 · score 0.59 · Mistral Small, system prompts, instruction, models, transcript, classification
  5. [5] § Methods › Models for classification › Zero-shot classification. ↔ zeroshot/zeroshotPatient.py, lines 1–37 · score 0.58 · Mistral Small, system prompts, instruction, models, transcript, classification
  6. [6] § Methods › Feature extraction › Semantic features › Embeddings and cosine similarity. ↔ semantic.ipynb, lines 39–70 · score 0.57 · sentence embedding, 1–3, split, SBERT, variance, window
  7. [7] § Methods › Data and preprocessing ↔ zeroshot/zeroshot.py, lines 51–92 · score 0.57 · important event, picture descriptions, cat, match, HCA, transcripts
  8. [8] § Methods › Data and preprocessing ↔ zeroshot/zeroshotMinimal.py, lines 47–89 · score 0.57 · important event, picture descriptions, cat, match, HCA, transcripts
  9. [9] § Methods › Feature extraction › Semantic features › Embeddings and cosine similarity. ↔ semantic.ipynb, lines 72–80 · score 0.56 · pre trained, Word embedding, fastText, vectors

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 394 lines · 14 KB · GPL-3.0 · 3 matches

  1. """Compute complexity metrics from Universal Dependencies.
  2. Usage: python3 udstyle.py [OPTIONS] FILE...
  3. --parse=LANG parse texts with Stanza; provide 2 letter language code
  4. --output=FILENAME write result to a tab-separated file.
  5. --persentence report per sentence results, not mean per document
  6. Reported metrics:
  7. - LEN: mean sentence length in words (excluding punctuation).
  8. - MDD: mean dependency distance (Gibson, 1998).
  9. - NDD: normalized dependency distance (Lei & Jockers, 2018).
  10. - ADJD: proportion of adjacent dependencies.
  11. - LEFT: dependency direction: proportion of left dependents.
  12. - MOD: nominal modifiers (Biber & Gray, 2010).
  13. - CLS: number of clauses per sentence.
  14. - CLL: average clause length (clauses/words)
  15. - LXD: lexical density: ratio of content words over total number of words
  16. - POS/DEP tag frequencies (only with --output)
  17. Example:
  18. $ python3 udstyle.py UD_Dutch-LassySmall/*.conllu
  19. LEN MDD NDD ADJD LEFT MOD CLS CLL LXD
  20. dev.conllu 14.182 2.461 0.926 0.500 0.459 0.052 2.223 9.190 0.603
  21. test.conllu 11.434 2.192 0.807 0.547 0.412 0.074 1.771 9.013 0.657
  22. train.conllu 11.027 2.172 0.775 0.564 0.391 0.072 1.863 8.107 0.645
  23. """
  24. import os
  25. import sys
  26. import getopt
  27. import subprocess
  28. from math import log, sqrt
  29. from contextlib import contextmanager
  30. from collections import Counter
  31. import pandas as pd
  32. # TODO: extract POS and syntactic n-grams frequencies
  33. # TODO: POS surprisal, requires training e.g. n-gram model on corpus
  34. # Constants for field numbers:
  35. ID, FORM, LEMMA, UPOS, XPOS, FEATS, HEAD, DEPREL, DEPS, MISC = range(10)
  36. # https://universaldependencies.org/format.html
  37. # ID: Word index, integer starting at 1 for each new sentence; may be a range
  38. # for multiword tokens; may be a decimal number for empty nodes (decimal
  39. # numbers can be lower than 1 but must be greater than 0).
  40. # FORM: Word form or punctuation symbol.
  41. # LEMMA: Lemma or stem of word form.
  42. # UPOS: Universal part-of-speech tag.
  43. # XPOS: Language-specific part-of-speech tag; underscore if not available.
  44. # FEATS: List of morphological features from the universal feature inventory or
  45. # from a defined language-specific extension; underscore if not
  46. # available.
  47. # HEAD: Head of the current word, which is either a value of ID or zero (0).
  48. # DEPREL: Universal dependency relation to the HEAD (root iff HEAD = 0) or a
  49. # defined language-specific subtype of one.
  50. # DEPS: Enhanced dependency graph in the form of a list of head-deprel pairs.
  51. # MISC: Any other annotation.
  52. def which(program, exception=True):
  53. """Return first match for program in search path.
  54. :param exception: By default, ValueError is raised when program not found.
  55. Pass False to return None in this case."""
  56. for path in os.environ.get('PATH', os.defpath).split(":"):
  57. if path and os.path.exists(os.path.join(path, program)):
  58. return os.path.join(path, program)
  59. if exception:
  60. raise ValueError('%r not found in path; please install it.' % program)
  61. @contextmanager
  62. def genericdecompressor(cmd, filename, encoding='utf8'):
  63. """Run command line decompressor on file and return file object.
  64. :param cmd: executable in path with gzip-like command line interface;
  65. e.g., ``gzip, zstd, lz4, bzip2, lzop``
  66. :param filename: the file to decompress.
  67. :param encoding: if None, mode is binary; otherwise, text.
  68. :raises ValueError: if command returns an error.
  69. :returns: a file-like object that must be used in a with-statement;
  70. supports .read() and iteration, but not seeking."""
  71. with subprocess.Popen(
  72. [which(cmd), '--decompress', '--stdout', '--quiet', filename],
  73. stdout=subprocess.PIPE, stderr=subprocess.PIPE,
  74. encoding=encoding) as proc:
  75. # FIXME: should use select to avoid deadlocks due to OS pipe buffers
  76. # filling up and blocking the child process.
  77. yield proc.stdout
  78. retcode = proc.wait()
  79. if retcode: # FIXME: retcode 2 means warning. allow warnings?
  80. raise ValueError('non-zero exit code %s from compressor %s:\n%r'
  81. % (retcode, cmd, proc.stderr.read()))
  82. def openread(filename, encoding='utf8'):
  83. """Open stdin/file for reading; decompress gz/lz4/zst files on-the-fly.
  84. :param encoding: if None, mode is binary; otherwise, text."""
  85. mode = 'rb' if encoding is None else 'rt'
  86. if filename == '-': # TODO: decompress stdin on-the-fly
  87. return open(sys.stdin.fileno(), mode=mode, encoding=encoding)
  88. if not isinstance(filename, int):
  89. if filename.endswith('.gz'):
  90. return genericdecompressor('gzip', filename, encoding)
  91. elif filename.endswith('.zst'):
  92. return genericdecompressor('zstd', filename, encoding)
  93. elif filename.endswith('.lz4'):
  94. return genericdecompressor('lz4', filename, encoding)
  95. return open(filename, mode=mode, encoding=encoding)
  96. def parsefiles(filenames, lang):
  97. """Parse UTF-8 encoded plain text files with Stanza if a corresponding
  98. .conllu file does not exist already."""
  99. nlp = None
  100. newfilenames = []
  101. for filename in filenames:
  102. conllu = '%s.conllu' % os.path.splitext(filename)[0]
  103. newfilenames.append(conllu)
  104. if (not os.path.exists(conllu)
  105. or os.stat(conllu).st_mtime < os.stat(filename).st_mtime):
  106. if nlp is None:
  107. import stanza
  108. from stanza.utils.conll import CoNLL
  109. try:
  110. nlp = stanza.Pipeline(lang)
  111. except FileNotFoundError:
  112. stanza.download(lang)
  113. nlp = stanza.Pipeline(lang)
  114. with open(filename, encoding='utf8') as inp:
  115. doc = nlp(inp.read())
  116. # TODO: preserve paragraph breaks
  117. CoNLL.write_doc2conll(doc, conllu)
  118. return newfilenames
  119. def conllureader(filename, excludepunct=False):
  120. """Load corpus. Returns list of lists of lists:
  121. sentences[sentno][tokenno][fieldno]"""
  122. result = []
  123. sent = []
  124. with openread(filename) as inp:
  125. for line in inp:
  126. if line == '\n':
  127. if sent:
  128. try:
  129. result.append(renumber(sent))
  130. except KeyError:
  131. pass
  132. sent = []
  133. elif line.startswith('#'): # ignore all comments
  134. pass
  135. else:
  136. fields = line[:-1].split('\t')
  137. if '.' in fields[ID]: # skip empty nodes
  138. continue
  139. elif excludepunct and fields[UPOS] == 'PUNCT':
  140. continue
  141. elif '-' in fields[ID]: # multiword tokens
  142. fields[ID] = int(fields[ID][:fields[ID].index('-')])
  143. else: # normal tokens
  144. fields[ID] = int(fields[ID])
  145. try:
  146. fields[HEAD] = int(fields[HEAD])
  147. except ValueError:
  148. continue
  149. sent.append(fields)
  150. if not result:
  151. raise ValueError('no sentences; not a valid .conllu file?')
  152. return result
  153. def renumber(sent):
  154. """Fix non-contiguous IDs because of multiword tokens or removed tokens"""
  155. mapping = {line[ID]: n for n, line in enumerate(sent, 1)}
  156. mapping[0] = 0
  157. for line in sent:
  158. line[ID] = mapping[line[ID]]
  159. line[HEAD] = mapping[line[HEAD]]
  160. return sent
  161. def mean(iterable):
  162. """Arithmetic mean."""
  163. seq = list(iterable) # accept generators
  164. return sum(seq) / len(seq)
  165. def analyze(filename, excludepunct=True, persentence=False):
  166. """Return a dict {featname: vector, ...} describing UD file.
  167. Each feature vector has a value for each sentence."""
  168. sentences = conllureader(filename, excludepunct=excludepunct)
  169. result = complexitymetrics(sentences)
  170. if persentence:
  171. result['sent'] = [' '.join(line[FORM] for line in sent)
  172. for sent in sentences]
  173. result['filename'] = filename
  174. return result
  175. # Get macro average over the per-sentence scores.
  176. # Might want to look at standard deviation and other aspects of the
  177. # distribution. TODO: offer micro average as well.
  178. for a, b in result.items():
  179. result[a] = sum(b) / len(b)
  180. result.update(counttags(sentences))
  181. return result
  182. def complexitymetrics(sentences):
  183. """Return dict of complexity metrics with results for each sentence."""
  184. result = {}
  185. result['LEN'] = [len(sent) for sent in sentences]
  186. # Ignore certain relations, following Chen and Gerdes (2017, p. 57)
  187. # http://www.aclweb.org/anthology/W17-6508
  188. exclude = ('fixed', 'flat', 'conj', 'punct')
  189. # Gibson (1998) http://dx.doi.org/10.1016/S0010-0277(98)00034-1
  190. # Liu (2008) https://hdl.handle.net/10371/70907
  191. # mean dependency distance
  192. result['MDD'] = [
  193. mean(abs(line[ID] - line[HEAD]) for line in sent
  194. if line[DEPREL] not in exclude)
  195. for sent in sentences]
  196. # Lei & Jockers (2018): https://doi.org/10.1080/09296174.2018.1504615
  197. # normalized dependency distance
  198. result['NDD'] = [
  199. abs(log(mdd
  200. / sqrt(
  201. ([line[DEPREL] for line in sent].index('root') + 1)
  202. * len(sent))))
  203. for mdd, sent in zip(result['MDD'], sentences)]
  204. # proportion of adjacent dependencies
  205. # https://doi.org/10.1016/j.langsci.2016.09.006
  206. result['ADJD'] = [mean(abs(line[ID] - line[HEAD]) == 1 for line in sent)
  207. for sent in sentences]
  208. # dependency direction: proportion of left dependents
  209. # http://www.aclweb.org/anthology/W17-6508
  210. result['LEFT'] = [mean(line[ID] < line[HEAD] for line in sent)
  211. for sent in sentences]
  212. # nominal modifiers;
  213. # attempt to measure phrasal complexity (as opposed to clausal complexity).
  214. # see e.g. https://doi.org/10.1016/j.jeap.2010.01.001
  215. result['MOD'] = [mean(line[DEPREL] == 'nmod' for line in sent)
  216. for sent in sentences]
  217. # number of clauses per sentence; https://doi.org/10.1007/s11145-007-9107-5
  218. result['CLS'] = [1 + sum(line[UPOS] == 'VERB' for line in sent)
  219. for sent in sentences]
  220. # avg clause len (clauses/words) https://aclanthology.org/2020.lrec-1.883
  221. result['CLL'] = [
  222. len(sent) / max(1, sum(line[UPOS] == 'VERB' for line in sent))
  223. for sent in sentences]
  224. # lexical density: ratio of content words over total number of words
  225. # https://aclanthology.org/2020.lrec-1.883
  226. content = ('ADJ', 'ADV', 'INTJ', 'NOUN', 'PROPN', 'VERB')
  227. result['LXD'] = [sum(line[UPOS] in content for line in sent) / len(sent)
  228. for sent in sentences]
  229. return result
  230. def counttags(sentences):
  231. """Count POS and dependency tags; returns relative frequencies."""
  232. numtokens = sum(len(sent) for sent in sentences)
  233. postags = Counter(line[UPOS] for sent in sentences for line in sent)
  234. deptags = Counter(line[DEPREL] for sent in sentences for line in sent)
  235. tags = {a: postags[a] / numtokens for a in [
  236. 'ADJ', # adjective
  237. 'ADP', # adposition
  238. 'ADV', # adverb
  239. 'AUX', # auxiliary
  240. 'CCONJ', # coordinating conjunction
  241. 'DET', # determiner
  242. 'INTJ', # interjection
  243. 'NOUN', # noun
  244. 'NUM', # numeral
  245. 'PART', # particle
  246. 'PRON', # pronoun
  247. 'PROPN', # proper noun
  248. 'PUNCT', # punctuation
  249. 'SCONJ', # subordinating conjunction
  250. 'SYM', # symbol
  251. 'VERB', # verb
  252. 'X', # other
  253. ]}
  254. tags.update({a: deptags[a] / numtokens for a in [
  255. 'acl', # clausal modifier of noun (adnominal clause)
  256. 'acl:relcl', # relative clause modifier
  257. 'advcl', # adverbial clause modifier
  258. 'advmod', # adverbial modifier
  259. 'advmod:emph', # emphasizing word, intensifier
  260. 'advmod:lmod', # locative adverbial modifier
  261. 'amod', # adjectival modifier
  262. 'appos', # appositional modifier
  263. 'aux', # auxiliary
  264. 'aux:pass', # passive auxiliary
  265. 'case', # case marking
  266. 'cc', # coordinating conjunction
  267. 'cc:preconj', # preconjunct
  268. 'ccomp', # clausal complement
  269. 'clf', # classifier
  270. 'compound', # compound
  271. 'compound:lvc', # light verb construction
  272. 'compound:prt', # phrasal verb particle
  273. 'compound:redup', # reduplicated compounds
  274. 'compound:svc', # serial verb compounds
  275. 'conj', # conjunct
  276. 'cop', # copula
  277. 'csubj', # clausal subject
  278. 'csubj:pass', # clausal passive subject
  279. 'dep', # unspecified dependency
  280. 'det', # determiner
  281. 'det:numgov', # pronominal quantifier governing the case of the noun
  282. 'det:nummod', # pronominal quantifier agreeing in case with the noun
  283. 'det:poss', # possessive determiner
  284. 'discourse', # discourse element
  285. 'dislocated', # dislocated elements
  286. 'expl', # expletive
  287. 'expl:impers', # impersonal expletive
  288. 'expl:pass', # reflexive pronoun used in reflexive passive
  289. 'expl:pv', # reflexive clitic with an inherently reflexive verb
  290. 'fixed', # fixed multiword expression
  291. 'flat', # flat multiword expression
  292. 'flat:foreign', # foreign words
  293. 'flat:name', # names
  294. 'goeswith', # goes with
  295. 'iobj', # indirect object
  296. 'list', # list
  297. 'mark', # marker
  298. 'nmod', # nominal modifier
  299. 'nmod:poss', # possessive nominal modifier
  300. 'nmod:tmod', # temporal modifier
  301. 'nsubj', # nominal subject
  302. 'nsubj:pass', # passive nominal subject
  303. 'nummod', # numeric modifier
  304. 'nummod:gov', # numeric modifier governing the case of the noun
  305. 'obj', # object
  306. 'obl', # oblique nominal
  307. 'obl:agent', # agent modifier
  308. 'obl:arg', # oblique argument
  309. 'obl:lmod', # locative modifier
  310. 'obl:tmod', # temporal modifier
  311. 'orphan', # orphan
  312. 'parataxis', # parataxis
  313. 'punct', # punctuation
  314. 'reparandum', # overridden disfluency
  315. 'root', # root
  316. 'vocative', # vocative
  317. 'xcomp', # open clausal complement
  318. ]})
  319. return tags
  320. def compare(filenames, parse=None, excludepunct=True, persentence=False):
  321. """Collect statistics for multiple files.
  322. Returns a dataframe with one row per filename, with the mean score
  323. for each metric in the colmuns."""
  324. if parse:
  325. filenames = parsefiles(filenames, parse)
  326. if persentence:
  327. return pd.concat([
  328. pd.DataFrame(
  329. analyze(filename, excludepunct=excludepunct,
  330. persentence=persentence))
  331. for filename in filenames],
  332. ignore_index=True)
  333. return pd.DataFrame({
  334. os.path.basename(filename):
  335. analyze(filename, excludepunct=excludepunct)
  336. for filename in filenames}).T
  337. def main():
  338. """CLI."""
  339. try:
  340. opts, args = getopt.gnu_getopt(
  341. sys.argv[1:], '', ['output=', 'parse=', 'persentence'])
  342. opts = dict(opts)
  343. except getopt.GetoptError:
  344. print(__doc__)
  345. return
  346. if not args:
  347. print(__doc__)
  348. return
  349. result = compare(
  350. args, opts.get('--parse'), persentence='--persentence' in opts)
  351. if '--persentence' in opts:
  352. if '--output' in opts:
  353. result.to_csv(opts.get('--output'), sep='\t')
  354. else:
  355. print(result)
  356. elif '--output' in opts:
  357. result.to_csv(opts.get('--output'), sep='\t')
  358. else:
  359. selection = 'LEN MDD NDD ADJD LEFT MOD CLS CLL LXD'.split()
  360. print(result.loc[:, selection].round(3)) # skip tags
  361. if __name__ == '__main__':
  362. main()

udstyle.py at commit d628427, under GPL-3.0 · at the source

Overview

  1. University of Groningen Center for Language and Cognition, Netherlands
Institutions: University of Groningen (Netherlands)
Journal: Schizophrenia research. Cognition, volume 45, article 100443
Dates: received 26 February 2026; accepted 7 May 2026; published online 26 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1016/j.scog.2026.100443 · PMID 42238842 · PMCID PMC13226953 · OpenAlex W7162451776
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: schizophrenia / psychosis (population), cognitive (subfield)
Methods: Machine learning, Connectivity
Keywords: Schizophrenia spectrum disorders, Wernicke’s aphasia, Language and thought, Machine learning, LLMs
Topic: Neurobiology of Language and Bilingualism (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: not cited yet (Europe PMC); 54 references in the paper

Abstract

Schizophrenia spectrum disorders (SSD) and Wernicke’s aphasia (WA) both disrupt meaningful speech, yet they arise from fundamentally different disturbances in thought and language. SSD is defined by formal thought disorder, in which disorganized thinking is inferred from abnormalities in speech, whereas WA reflects a primary breakdown of language implementation following focal brain damage. We investigated whether quantitative markers of lexical–semantic, syntactic structure, and semantic coherence in spontaneous speech can distinguish SSD, WA, and healthy controls.

Using Natural Language Processing techniques, we extracted syntactic, lexical and local semantic similarity features from spontaneous speech transcripts and used them in supervised machine learning models to classify diagnostic groups. In parallel, an instruction-tuned large language model (LLM) was used in a zero-shot setting to assign transcripts to diagnostic categories and to track the severity of language disturbance.

Our results showed a distinct linguistic pattern, particularly in syntactic and local semantic organization for WA, indicating a paradigmatic language disorder. By contrast, the same features were less effective in distinguishing SSD from matched controls, in line with the view that SSD reflects a more diffuse disturbance of thought that only partially manifests in surface language. Zero-shot LLM classifications approached the performance of supervised models for WA-related contrasts and were sensitive to graded language disturbance. At the same time, strong task and dataset effects underscored the need for carefully controlled speech elicitation. Together, these findings highlight both the promise and the limitations of automated language analysis for clinical diagnostics and for understanding speech and thought abnormalities.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.

P-Zande/ssd-vs-wa

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: a5cb25cf01f9733ec807ff019e2c47064df5baa8, 26 January 2026
Languages: Python (5), Jupyter (4)
Size: 11 files, 9 scripts
Software Heritage: not archived
Found in: the text
Holds: README, environment (requirements.txt), 4 notebooks
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: pandas (9 files), Hugging Face Transformers (5 files), PyTorch (4 files), NumPy (3 files), scikit-learn (3 files)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
10 files

andreasvc/udstyle

License: GPL-3.0
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: d628427dc0b3bb36872b48411f7bb7c59acdfd62, 18 June 2026
Languages: Python (1)
Size: 4 files, 1 script
Software Heritage: archived
Found in: the references
Holds: README, license file, CITATION.cff
Not found: environment file, tests, continuous integration, documentation
Tools: pandas (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
3 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 10 scripts, each with its path and the digest of its content;
  • 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Authors: added Perry van der Zande (0009-0000-2004-8341); removed Perry van der Zande

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 3 authors, 5 keywords, 41 references.

Cite

This paper

van der Zande, P., van Cranenburgh, A., & Tsiwah, F. (2026). From thought to language: Comparing schizophrenia spectrum disorders and Wernicke's aphasia with machine learning and LLMs. Schizophrenia research. Cognition, 45, 100443. https://doi.org/10.1016/j.scog.2026.100443

BibTeX

@article{vanderzande2026thought,
author = {van der Zande, Perry and van Cranenburgh, Andreas and Tsiwah, Frank},
title = {{From thought to language: Comparing schizophrenia spectrum disorders and Wernicke's aphasia with machine learning and LLMs}},
journal = {Schizophrenia research. Cognition},
year = {2026},
month = may,
volume = {45},
pages = {100443},
publisher = {Elsevier},
issn = {2215-0013},
doi = {10.1016/j.scog.2026.100443},
url = {https://doi.org/10.1016/j.scog.2026.100443},
pmid = {42238842},
pmcid = {PMC13226953}
}

RIS

TY - JOUR
AU - van der Zande, Perry
AU - van Cranenburgh, Andreas
AU - Tsiwah, Frank
TI - From thought to language: Comparing schizophrenia spectrum disorders and Wernicke's aphasia with machine learning and LLMs
T2 - Schizophrenia research. Cognition
J2 - Schizophr Res Cogn
PY - 2026
DA - 2026/05/26
VL - 45
SP - 100443
SN - 2215-0013
PB - Elsevier
DO - 10.1016/j.scog.2026.100443
UR - https://doi.org/10.1016/j.scog.2026.100443
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.scog.2026.100443",
"type": "article-journal",
"title": "From thought to language: Comparing schizophrenia spectrum disorders and Wernicke's aphasia with machine learning and LLMs",
"container-title": "Schizophrenia research. Cognition",
"author": [
{
"family": "van der Zande",
"given": "Perry"
},
{
"family": "van Cranenburgh",
"given": "Andreas"
},
{
"family": "Tsiwah",
"given": "Frank"
}
],
"container-title-short": "Schizophr Res Cogn",
"volume": "45",
"page": "100443",
"DOI": "10.1016/j.scog.2026.100443",
"PMID": "42238842",
"PMCID": "PMC13226953",
"ISSN": "2215-0013",
"publisher": "Elsevier",
"URL": "https://doi.org/10.1016/j.scog.2026.100443",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
26
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-75841-9 [code]
Model-based semantic distance reveals adaptive coordination of distinct cognitive systems in flexible knowledge retrieval.
Journal: Nature communications
In common: Hugging Face Transformers, PyTorch, pandas, 1 other tool, cognitive, 2 references
[2] doi:10.7554/elife.107933 [code]
Modality-agnostic decoding of vision and language from fMRI.
Journal: eLife
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, cognitive, 1 reference
[3] doi:10.1016/j.isci.2026.116135 [code]
Predicting ICU in-hospital mortality from text-encoded structured EHR data using adaptive transformer layer fusion.
Journal: iScience
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, 1 reference
[4] doi:10.1038/s41467-026-74358-5 [code]
Brain-inspired spatial intelligence for embodied agents.
Journal: Nature communications
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, cognitive
[5] doi:10.1038/s42003-026-10011-7 [code]
Learning brain dynamics across distinct scaling regimes reveals psychiatric signatures.
Journal: Communications biology
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, cognitive
[6] doi:10.1038/s42003-026-10169-0 [code]
Shared representations in brains and models reveal a two-route cortical organization during scene perception.
Journal: Communications biology
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, cognitive
[7] doi:10.1162/imag.a.1227 [code]
Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, cognitive
[8] doi:10.7554/elife.106543 [code]
Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs.
Journal: eLife
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, cognitive
[9] doi:10.1038/s41467-026-71267-5 [code]
Human-like cognitive generalization for large models via mental representation-guided supervision.
Journal: Nature communications
In common: Hugging Face Transformers, PyTorch, scikit-learn, 2 other tools, cognitive
[10] doi:10.1162/nol.a.244 [code]
A Novel Approach to Map the Causal Impact of Brain Stimulation on Semantic Processing With Language Models.
Journal: Neurobiology of language (Cambridge, Mass.)
In common: Hugging Face Transformers, PyTorch, scikit-learn, 1 other tool, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.