OSCR

Benchmarking a Local Schema-Constrained Large Language Model Pipeline for Abstract Screening and Evidence Mapping.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § Materials and methods ↔ scripts/pubmed1_screen.py, lines 1–31 · score 0.62 · gpt oss, logs, Ollama, Max, append, JSONL
  2. [2] § Materials and methods ↔ scripts/pubmed_prisma.py, lines 99–138 · score 0.59 · reason_code, PubMed, PRISMA, ineligible, workflow, human

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 628 lines · 21 KB · MIT · 1 match

  1. #!/usr/bin/env python3
  2. """LLM-assisted title/abstract screening for remaining 'maybe' records.
  3. This script:
  4. - Reads all runs/<query_id>/records.json (produced by pubmed1.py)
  5. - Builds the post-automation screening universe (unique PMIDs)
  6. - Reads runs/decisions.jsonl (append-only audit log)
  7. - Selects PMIDs whose latest decision at a stage is missing or 'maybe'
  8. - Calls Ollama locally (ollama run <model>) to classify each record
  9. - Appends new decision lines to decisions.jsonl with by='LLM'
  10. Design goals: simple, local, reproducible, low-hallucination.
  11. Usage:
  12. python3 pubmed1_screen.py
  13. python3 pubmed1_screen.py runs --max 20 --dry-run
  14. python3 pubmed1_screen.py runs --model gpt-oss:20b
  15. """
  16. from __future__ import annotations
  17. import argparse
  18. import json
  19. import re
  20. import subprocess
  21. import requests
  22. import ast
  23. import textwrap
  24. from datetime import datetime, timezone
  25. from pathlib import Path
  26. from typing import Any, Dict, Iterable, List, Optional, Set, Tuple
  27. # Keep these aligned with pubmed1prisma.py automation defaults
  28. DEFAULT_AUTOMATION_INELIGIBLE_PUBTYPES = {
  29. "editorial",
  30. "letter",
  31. "comment",
  32. "news",
  33. "interview",
  34. "published erratum",
  35. "retracted publication",
  36. "retraction of publication",
  37. }
  38. # Screening reason codes: keep small and PRISMA-friendly.
  39. EXCLUSION_REASON_CODES = {
  40. "NOT_RELEVANT",
  41. "WRONG_DESIGN",
  42. "WRONG_POPULATION",
  43. "WRONG_OUTCOME",
  44. "WRONG_INTERVENTION",
  45. "NO_ABSTRACT",
  46. "LANGUAGE",
  47. "OTHER",
  48. }
  49. INCLUSION_REASON_CODES = {"ELIGIBLE", "POTENTIALLY_ELIGIBLE"}
  50. def _utc_now_iso() -> str:
  51. return datetime.now(timezone.utc).replace(microsecond=0).isoformat()
  52. def _norm_str(x: Any) -> str:
  53. return str(x).strip().lower()
  54. def _read_criteria_text(criteria_path: Path) -> str:
  55. """Read criteria text from a file and return it as plain text.
  56. This is the authoritative criteria block inserted into the LLM prompt.
  57. """
  58. if not criteria_path.exists():
  59. raise FileNotFoundError(f"Criteria file not found: {criteria_path}")
  60. txt = criteria_path.read_text(encoding="utf-8").strip()
  61. if not txt:
  62. raise ValueError(f"Criteria file is empty: {criteria_path}")
  63. return txt + "\n"
  64. def _load_records_json(path: Path) -> List[dict[str, Any]]:
  65. with path.open("r", encoding="utf-8") as f:
  66. data = json.load(f)
  67. if not isinstance(data, dict):
  68. raise ValueError(f"{path}: expected a JSON object")
  69. records = data.get("records")
  70. if not isinstance(records, list):
  71. raise ValueError(f"{path}: missing or invalid 'records' list")
  72. return [r for r in records if isinstance(r, dict)]
  73. def _find_records_files(runs_dir: Path) -> List[Path]:
  74. if not runs_dir.exists():
  75. raise FileNotFoundError(f"Runs directory not found: {runs_dir}")
  76. if not runs_dir.is_dir():
  77. raise NotADirectoryError(f"Not a directory: {runs_dir}")
  78. files = sorted(runs_dir.rglob("records.json"))
  79. if not files:
  80. raise FileNotFoundError(
  81. f"No records.json found under {runs_dir}. Expected runs/<query_id>/records.json"
  82. )
  83. return files
  84. def _iter_records(runs_dir: Path) -> Iterable[dict[str, Any]]:
  85. for fp in _find_records_files(runs_dir):
  86. for r in _load_records_json(fp):
  87. yield r
  88. def _automation_ineligible(r: dict[str, Any], *, exclude_nonhuman: bool = True, exclude_pubtypes: bool = True) -> bool:
  89. if exclude_nonhuman:
  90. is_humans = r.get("is_humans")
  91. if is_humans is False:
  92. return True
  93. if exclude_pubtypes:
  94. pts = r.get("publication_types") or []
  95. pts_norm = {_norm_str(p) for p in pts}
  96. if pts_norm & DEFAULT_AUTOMATION_INELIGIBLE_PUBTYPES:
  97. return True
  98. return False
  99. def _read_decisions_jsonl(path: Path) -> List[dict[str, Any]]:
  100. if not path.exists():
  101. return []
  102. out: List[dict[str, Any]] = []
  103. with path.open("r", encoding="utf-8") as f:
  104. for line_no, line in enumerate(f, start=1):
  105. line = line.strip()
  106. if not line:
  107. continue
  108. try:
  109. obj = json.loads(line)
  110. except json.JSONDecodeError:
  111. raise ValueError(f"{path}:{line_no}: invalid JSON (not JSONL?)")
  112. if isinstance(obj, dict):
  113. out.append(obj)
  114. return out
  115. def _latest_decision_by_pmid_stage(decisions: List[dict[str, Any]]) -> dict[Tuple[str, str], dict[str, Any]]:
  116. latest: dict[Tuple[str, str], dict[str, Any]] = {}
  117. latest_ts: dict[Tuple[str, str], datetime] = {}
  118. for obj in decisions:
  119. pmid = obj.get("pmid")
  120. stage = obj.get("stage")
  121. if not pmid or not stage:
  122. continue
  123. key = (str(pmid), str(stage))
  124. ts_raw = obj.get("timestamp")
  125. ts = None
  126. if isinstance(ts_raw, str):
  127. try:
  128. ts = datetime.fromisoformat(ts_raw.replace("Z", "+00:00"))
  129. except Exception:
  130. ts = None
  131. if ts is None:
  132. latest[key] = obj
  133. continue
  134. prev = latest_ts.get(key)
  135. if prev is None or ts >= prev:
  136. latest_ts[key] = ts
  137. latest[key] = obj
  138. return latest
  139. def _ensure_trailing_newline(path: Path) -> None:
  140. """Ensure file ends with newline so JSONL append can't corrupt the last line."""
  141. if not path.exists() or path.stat().st_size == 0:
  142. return
  143. with path.open("rb") as fb:
  144. fb.seek(-1, 2)
  145. last = fb.read(1)
  146. if last != b"\n":
  147. with path.open("a", encoding="utf-8") as f:
  148. f.write("\n")
  149. # Helper: Write raw model output for debugging failed parses
  150. def _write_debug_raw(debug_dir: Path, pmid: str, tag: str, content: str) -> None:
  151. debug_dir.mkdir(parents=True, exist_ok=True)
  152. safe_pmid = re.sub(r"[^0-9A-Za-z._-]+", "_", pmid)
  153. ts = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
  154. out = debug_dir / f"{safe_pmid}.{tag}.{ts}.txt"
  155. out.write_text(content or "", encoding="utf-8")
  156. def _extract_first_json_object(text: str) -> Optional[dict[str, Any]]:
  157. """Extract the first JSON object from model output.
  158. We ask for JSON-only output, but models sometimes wrap it in prose, code fences,
  159. or emit Python dicts / trailing commas. We keep fixes conservative.
  160. """
  161. if not text:
  162. return None
  163. s = text.strip()
  164. # Fast path: whole output is a JSON object
  165. if s.startswith("{") and s.endswith("}"):
  166. try:
  167. obj = json.loads(s)
  168. return obj if isinstance(obj, dict) else None
  169. except Exception:
  170. pass
  171. # Find first balanced {...} block (more reliable than greedy regex)
  172. start = s.find("{")
  173. if start < 0:
  174. return None
  175. depth = 0
  176. end = None
  177. for i, ch in enumerate(s[start:], start=start):
  178. if ch == "{":
  179. depth += 1
  180. elif ch == "}":
  181. depth -= 1
  182. if depth == 0:
  183. end = i + 1
  184. break
  185. if end is None:
  186. return None
  187. chunk = s[start:end].strip()
  188. # Attempt strict JSON
  189. try:
  190. obj = json.loads(chunk)
  191. return obj if isinstance(obj, dict) else None
  192. except Exception:
  193. pass
  194. # Conservative normalization
  195. norm = chunk
  196. norm = norm.replace("\u201c", '"').replace("\u201d", '"').replace("\u2018", "'").replace("\u2019", "'")
  197. # Remove trailing commas before } or ]
  198. norm = re.sub(r",\s*(\}|\])", r"\1", norm)
  199. # Quote common bareword enum mistakes (keep scope narrow)
  200. # e.g. "reason_code": WRONG_INTERVENTION -> "reason_code": "WRONG_INTERVENTION"
  201. norm = re.sub(
  202. r'("reason_code"\s*:\s*)([A-Za-z_][A-Za-z0-9_]*)(\s*[\},])',
  203. r'\1"\2"\3',
  204. norm,
  205. )
  206. norm = re.sub(
  207. r'("decision"\s*:\s*)(include|exclude|maybe)(\s*[\},])',
  208. r'\1"\2"\3',
  209. norm,
  210. flags=re.IGNORECASE,
  211. )
  212. norm = re.sub(
  213. r'("confidence"\s*:\s*)(low|medium|high)(\s*[\},])',
  214. r'\1"\2"\3',
  215. norm,
  216. flags=re.IGNORECASE,
  217. )
  218. try:
  219. obj = json.loads(norm)
  220. return obj if isinstance(obj, dict) else None
  221. except Exception:
  222. pass
  223. # Last resort: model may have emitted a Python dict
  224. try:
  225. obj2 = ast.literal_eval(norm)
  226. return obj2 if isinstance(obj2, dict) else None
  227. except Exception:
  228. return None
  229. def _build_prompt(record: dict[str, Any], criteria_text: str) -> str:
  230. pmid = str(record.get("pmid") or "")
  231. title = (record.get("title") or "").strip()
  232. abstract = (record.get("abstract") or "").strip()
  233. pubtypes = record.get("publication_types") or []
  234. languages = record.get("languages") or []
  235. is_humans = record.get("is_humans")
  236. return (
  237. "You are screening PubMed records for a systematic review with STRICT eligibility criteria.\n"
  238. "Use ONLY the provided Title/Abstract/metadata and the AUTHORITATIVE CRITERIA below.\n"
  239. "Do NOT infer missing details.\n"
  240. "If key information is missing, set decision=\"maybe\" and reason_code=\"UNCLEAR\".\n\n"
  241. "AUTHORITATIVE SCREENING CRITERIA (apply strictly):\n"
  242. + criteria_text
  243. + "\n"
  244. "Return EXACTLY one JSON object (no markdown, no commentary) with keys:\n"
  245. "Output MUST start with '{' and end with '}' and contain nothing else (single JSON object only).\n"
  246. "All string values MUST be in double quotes, including reason_code (e.g., \"reason_code\": \"WRONG_INTERVENTION\").\n"
  247. "pmid, decision, reason_code, confidence, notes\n"
  248. "Where decision is one of: include, exclude, maybe.\n"
  249. "reason_code MUST be consistent with decision:\n"
  250. "- if decision=exclude: choose one of NOT_RELEVANT, WRONG_DESIGN, WRONG_POPULATION, WRONG_OUTCOME, WRONG_INTERVENTION, NO_ABSTRACT, LANGUAGE, OTHER\n"
  251. "- if decision=include: use ELIGIBLE\n"
  252. "- if decision=maybe: use POTENTIALLY_ELIGIBLE or UNCLEAR\n"
  253. "confidence: low|medium|high\n"
  254. "notes: <= 20 words, strictly quoting or paraphrasing title/abstract (no speculation).\n\n"
  255. "Decision rules:\n"
  256. "- INCLUDE only if ALL eligibility criteria are explicitly satisfied.\n"
  257. "- EXCLUDE with WRONG_DESIGN if review/meta/protocol/editorial.\n"
  258. "- If uncertain due to missing details, choose MAYBE with UNCLEAR.\n\n"
  259. f"PMID: {pmid}\n"
  260. f"PublicationTypes: {pubtypes}\n"
  261. f"Languages: {languages}\n"
  262. f"HumansFlag: {is_humans}\n\n"
  263. f"Title: {title}\n\n"
  264. f"Abstract: {abstract if abstract else 'NO ABSTRACT'}\n"
  265. )
  266. def _call_ollama_http(model: str, prompt: str, timeout_s: int = 180) -> str:
  267. """Call Ollama via local HTTP API and return the generated text.
  268. Uses POST http://localhost:11434/api/generate
  269. """
  270. url = "http://localhost:11434/api/generate"
  271. payload: Dict[str, Any] = {
  272. "model": model,
  273. "prompt": prompt,
  274. "stream": False,
  275. }
  276. try:
  277. r = requests.post(url, json=payload, timeout=timeout_s)
  278. except requests.RequestException as e:
  279. raise RuntimeError(f"Ollama HTTP request failed: {e}")
  280. if r.status_code != 200:
  281. # Ollama typically returns JSON with an 'error' field on failures.
  282. try:
  283. j = r.json()
  284. err = j.get("error") if isinstance(j, dict) else None
  285. except Exception:
  286. err = None
  287. msg = err or (r.text or "").strip() or f"HTTP {r.status_code}"
  288. raise RuntimeError(f"Ollama HTTP error: {msg}")
  289. try:
  290. data = r.json()
  291. except Exception:
  292. raise RuntimeError("Ollama HTTP response was not valid JSON")
  293. if not isinstance(data, dict):
  294. raise RuntimeError("Ollama HTTP response JSON had unexpected shape")
  295. out = data.get("response")
  296. if not isinstance(out, str):
  297. raise RuntimeError("Ollama HTTP response missing 'response' text")
  298. return out
  299. def _call_ollama_subprocess(model: str, prompt: str, timeout_s: int = 180) -> str:
  300. """Call Ollama via CLI (subprocess) and return raw stdout.
  301. We pass the prompt as a CLI argument (not stdin) to avoid interactive/stdin quirks.
  302. """
  303. proc = subprocess.run(
  304. ["ollama", "run", model, prompt],
  305. text=True,
  306. capture_output=True,
  307. timeout=timeout_s,
  308. )
  309. if proc.returncode != 0:
  310. err = (proc.stderr or "").strip()
  311. raise RuntimeError(f"ollama run failed (code {proc.returncode}): {err}")
  312. return proc.stdout
  313. def _call_ollama(model: str, prompt: str, timeout_s: int = 180, transport: str = "http") -> str:
  314. """Call Ollama using the selected transport ('http' or 'subprocess')."""
  315. t = (transport or "http").strip().lower()
  316. if t == "subprocess":
  317. return _call_ollama_subprocess(model, prompt, timeout_s=timeout_s)
  318. return _call_ollama_http(model, prompt, timeout_s=timeout_s)
  319. def _make_decision_line(pmid: str, stage: str, obj: dict[str, Any], by: str) -> Dict[str, Any]:
  320. decision = str(obj.get("decision") or "maybe").lower().strip()
  321. if decision not in {"include", "exclude", "maybe"}:
  322. decision = "maybe"
  323. rc = obj.get("reason_code")
  324. reason_code = str(rc).strip().upper() if rc is not None and str(rc).strip() else None
  325. # Normalize confidence
  326. conf = str(obj.get("confidence") or "").lower().strip()
  327. if conf not in {"low", "medium", "high"}:
  328. conf = None
  329. # Enforce decision/reason consistency (guard against model nonsense)
  330. if decision == "include":
  331. # If model provided an exclusion reason, flip to exclude.
  332. if reason_code in EXCLUSION_REASON_CODES or reason_code == "UNCLEAR":
  333. decision = "exclude"
  334. else:
  335. reason_code = "ELIGIBLE"
  336. if decision == "exclude":
  337. if reason_code not in EXCLUSION_REASON_CODES:
  338. # Default to OTHER if missing or invalid
  339. reason_code = "OTHER"
  340. if decision == "maybe":
  341. if reason_code is None:
  342. reason_code = "UNCLEAR"
  343. elif reason_code in EXCLUSION_REASON_CODES:
  344. # If they gave a clear exclusion reason, treat it as exclude.
  345. decision = "exclude"
  346. elif reason_code not in {"POTENTIALLY_ELIGIBLE", "UNCLEAR"}:
  347. reason_code = "UNCLEAR"
  348. # If the model outputs a different PMID, ignore it (we trust our input PMID).
  349. _pmid_out = obj.get("pmid")
  350. if _pmid_out and str(_pmid_out).strip() != pmid:
  351. # keep going; do not overwrite
  352. pass
  353. # Keep optional short notes
  354. notes = obj.get("notes")
  355. if notes is not None:
  356. notes = str(notes).strip()
  357. if len(notes) > 200:
  358. notes = notes[:200]
  359. line: Dict[str, Any] = {
  360. "pmid": pmid,
  361. "stage": stage,
  362. "decision": decision,
  363. "reason_code": reason_code,
  364. "by": by,
  365. "confidence": conf,
  366. "timestamp": _utc_now_iso(),
  367. }
  368. if notes:
  369. line["notes"] = notes
  370. return line
  371. def main() -> int:
  372. ap = argparse.ArgumentParser(description="LLM-assisted screening for remaining 'maybe' records")
  373. ap.add_argument(
  374. "runs_dir",
  375. nargs="?",
  376. default="runs",
  377. help="Path to runs/ folder (default: ./runs)",
  378. )
  379. ap.add_argument(
  380. "--decisions",
  381. default=None,
  382. help="Decisions JSONL path (default: <runs_dir>/decisions.jsonl)",
  383. )
  384. ap.add_argument(
  385. "--stage",
  386. default="abstract",
  387. choices=["abstract", "fulltext"],
  388. help="Decision stage to write (default: abstract)",
  389. )
  390. ap.add_argument(
  391. "--criteria",
  392. default="criteriaImpuls.txt",
  393. help="Path to criteria text file to inject into the prompt (default: criteriaImpuls.txt)",
  394. )
  395. ap.add_argument(
  396. "--model",
  397. default="gpt-oss:20b",
  398. help="Ollama model name (default: gpt-oss:20b)",
  399. )
  400. ap.add_argument(
  401. "--transport",
  402. default="http",
  403. choices=["http", "subprocess"],
  404. help="How to call Ollama (default: http). Use 'subprocess' for CLI fallback.",
  405. )
  406. ap.add_argument(
  407. "--by",
  408. default="LLM",
  409. help="Value for the 'by' field (default: LLM)",
  410. )
  411. ap.add_argument(
  412. "--max",
  413. type=int,
  414. default=0,
  415. help="Max number of PMIDs to process (0 = no limit)",
  416. )
  417. ap.add_argument(
  418. "--dry-run",
  419. action="store_true",
  420. default=False,
  421. help="Do not write anything; just show what would be processed",
  422. )
  423. ap.add_argument(
  424. "--timeout",
  425. type=int,
  426. default=180,
  427. help="Timeout seconds per record (default: 180)",
  428. )
  429. ap.add_argument(
  430. "--debug-raw-dir",
  431. default=None,
  432. help="If set, write raw LLM outputs for any parse failures to this directory (default: <runs_dir>/llm_raw)",
  433. )
  434. args = ap.parse_args()
  435. runs_dir = Path(args.runs_dir).expanduser().resolve()
  436. decisions_path = (
  437. Path(args.decisions).expanduser().resolve()
  438. if args.decisions
  439. else (runs_dir / "decisions.jsonl").resolve()
  440. )
  441. # Resolve criteria path (relative paths: script dir first, then current working dir)
  442. criteria_arg = Path(args.criteria).expanduser()
  443. if criteria_arg.is_absolute():
  444. criteria_path = criteria_arg
  445. else:
  446. script_dir = Path(__file__).resolve().parent
  447. candidate1 = (script_dir / criteria_arg).resolve()
  448. candidate2 = (Path.cwd() / criteria_arg).resolve()
  449. criteria_path = candidate1 if candidate1.exists() else candidate2
  450. criteria_text = _read_criteria_text(criteria_path)
  451. debug_raw_dir = (
  452. Path(args.debug_raw_dir).expanduser().resolve()
  453. if args.debug_raw_dir
  454. else (runs_dir / "llm_raw").resolve()
  455. )
  456. # Build screening universe (unique PMIDs post automation) and keep one representative record per PMID
  457. pmid_to_record: dict[str, dict[str, Any]] = {}
  458. for r in _iter_records(runs_dir):
  459. pmid = r.get("pmid")
  460. if not pmid:
  461. continue
  462. pmid_s = str(pmid)
  463. if pmid_s in pmid_to_record:
  464. continue
  465. if _automation_ineligible(r):
  466. continue
  467. pmid_to_record[pmid_s] = r
  468. decisions = _read_decisions_jsonl(decisions_path)
  469. latest = _latest_decision_by_pmid_stage(decisions)
  470. # Select candidates: missing or maybe
  471. candidates: List[str] = []
  472. for pmid in sorted(pmid_to_record.keys()):
  473. obj = latest.get((pmid, args.stage))
  474. current_dec = str((obj or {}).get("decision") or "missing").lower()
  475. if current_dec in {"missing", "maybe"}:
  476. candidates.append(pmid)
  477. if args.max and args.max > 0:
  478. candidates = candidates[: args.max]
  479. print(f"Runs dir: {runs_dir}")
  480. print(f"Decisions file: {decisions_path}")
  481. print(f"Screening universe (unique PMIDs, post-automation): {len(pmid_to_record)}")
  482. print(f"Candidates to send to LLM (missing/maybe): {len(candidates)}")
  483. print(f"Model: {args.model}")
  484. print(f"Transport: {args.transport}")
  485. print(f"Criteria file: {criteria_path}")
  486. if args.dry_run:
  487. if candidates:
  488. print("DRY-RUN PMIDs:")
  489. for pmid in candidates:
  490. print(f" - {pmid}")
  491. return 0
  492. if not candidates:
  493. print("Nothing to do.")
  494. return 0
  495. decisions_path.parent.mkdir(parents=True, exist_ok=True)
  496. _ensure_trailing_newline(decisions_path)
  497. appended = 0
  498. errors = 0
  499. for pmid in candidates:
  500. r = pmid_to_record[pmid]
  501. prompt = _build_prompt(r, criteria_text)
  502. try:
  503. raw = _call_ollama(args.model, prompt, timeout_s=args.timeout, transport=args.transport)
  504. obj = _extract_first_json_object(raw)
  505. raw2 = ""
  506. if not obj:
  507. # One retry with a terse "repair" instruction
  508. repair = (
  509. prompt
  510. + "\n\nIMPORTANT: Your previous output was not valid JSON. "
  511. + "Return ONLY one valid JSON object, starting with '{' and ending with '}', no extra text."
  512. )
  513. raw2 = _call_ollama(args.model, repair, timeout_s=args.timeout, transport=args.transport)
  514. obj = _extract_first_json_object(raw2)
  515. if not obj:
  516. _write_debug_raw(debug_raw_dir, pmid, "raw", raw)
  517. if raw2:
  518. _write_debug_raw(debug_raw_dir, pmid, "raw_retry", raw2)
  519. raise ValueError(
  520. f"Could not parse JSON from model output (saved raw to {debug_raw_dir})"
  521. )
  522. line = _make_decision_line(pmid, args.stage, obj, args.by)
  523. with decisions_path.open("a", encoding="utf-8") as f:
  524. f.write(json.dumps(line, ensure_ascii=False) + "\n")
  525. appended += 1
  526. print(f"OK {pmid}: {line['decision']} ({line.get('reason_code')})")
  527. except Exception as e:
  528. errors += 1
  529. print(f"ERROR {pmid}: {e}")
  530. print(" Hint: try running this PMID with --max 1 and/or increase --timeout; also ensure the model outputs JSON.")
  531. print(f"DONE. Appended: {appended}. Errors: {errors}.")
  532. return 0
  533. if __name__ == "__main__":
  534. raise SystemExit(main())

pubmed1_screen.py at commit b9e00a7, under MIT · at the source

Overview

Authors: Alessandro Serretti1,2
  1. Psychiatry, Università degli Studi di Enna Kore, Enna, ITA
  2. Psychiatry, Oasi Research Institute-IRCCS, Troina, ITA
Journal: Cureus, volume 18, issue 6, article e111193
Dates: accepted 20 June 2026; published online 20 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.7759/cureus.111193 · PMID 42483140 · PMCID PMC13384908 · OpenAlex W7165359315
Open access: diamond, a free copy (OpenAlex)
Status: code verified
Categories: human (organism)
Methods: fMRI & imaging, Machine learning
Keywords: brain imaging, large language models, mapping evidence, psychiatry, scoping review
Journal subjects: Psychiatry
Topic: Topic Modeling (Artificial Intelligence, Computer Science), according to OpenAlex
Citations: not cited yet (Europe PMC); 26 references in the paper

Abstract

Background

The growth of biomedical literature increasingly exceeds the capacity of manual evidence synthesis. Large language models (LLMs) may support abstract screening and structured extraction, but many current workflows depend on proprietary cloud APIs, creating challenges for governance, reproducibility, and scalable deployment.

Methods

I developed a fully local, open-weight, schema-constrained pipeline (gpt-oss-20b, deployed via Ollama on Apple M1 Max) for title/abstract-based scoping workflows. The pipeline combined deterministic metadata filtering, LLM-assisted screening, and structured abstract extraction. Performance was benchmarked against three published systematic reviews (ketamine/neuroimaging; clozapine/suicidality; clozapine patient/caregiver perspectives) using precision, recall, and F1 against reference inclusion sets. I also report audit-adjusted estimates (i.e., performance metrics recalculated after manual full-text adjudication of discrepant records) alongside standard reference-set performance.

Results

In the ketamine/neuroimaging benchmark, the pipeline retained all 41 studies included in the original review; after audit adjustment, recall was 100.0% (46/46), accuracy 99.4% (156/157), precision 97.9% (46/47), and F1 98.9%. For clozapine/suicidality, recall was 79.3% (46/58), and F1 was 76.0%, with missed studies largely attributable to missing or non-informative abstracts. For clozapine patient/caregiver perspectives, recall was 88.9% (56/63), and F1 was 83.6%, with similar abstract-level constraints. Abstract-level extraction recovered audited metadata fields without detected errors and generated evidence maps that were thematically concordant with the main narrative structure of the reference reviews.

Conclusions

As a proof-of-concept, a fully local LLM pipeline can support scalable and auditable abstract-based scoping and high-level evidence mapping. Because performance was benchmarked against three reviews with partly audit-adjusted reference sets, the findings require confirmation in larger, independently adjudicated evaluations. Random human audit remains advisable, and expert full-text synthesis remains necessary when abstracts are non-informative or when mechanistic precision is required.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

svonte/Scoping-Systematic

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: b9e00a784ad01794111f110b7699aadaf599fadd, 10 June 2026
Languages: Python (10)
Size: 21 files, 10 scripts
Software Heritage: not archived
Found in: the text, “Materials and methods”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: pandas (4 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
12 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 10 scripts, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 1 author, 5 keywords, 20 references.

Cite

This paper

Serretti, A. (2026). Benchmarking a Local Schema-Constrained Large Language Model Pipeline for Abstract Screening and Evidence Mapping. Cureus, 18(6), e111193. https://doi.org/10.7759/cureus.111193

BibTeX

@article{serretti2026benchmarking,
author = {Serretti, Alessandro},
title = {{Benchmarking a Local Schema-Constrained Large Language Model Pipeline for Abstract Screening and Evidence Mapping}},
journal = {Cureus},
year = {2026},
month = jun,
volume = {18},
number = {6},
pages = {e111193},
publisher = {Cureus Inc.},
issn = {2168-8184},
doi = {10.7759/cureus.111193},
url = {https://doi.org/10.7759/cureus.111193},
pmid = {42483140},
pmcid = {PMC13384908}
}

RIS

TY - JOUR
AU - Serretti, Alessandro
TI - Benchmarking a Local Schema-Constrained Large Language Model Pipeline for Abstract Screening and Evidence Mapping
T2 - Cureus
J2 - Cureus
PY - 2026
DA - 2026/06/20
VL - 18
IS - 6
SP - e111193
SN - 2168-8184
PB - Cureus Inc.
DO - 10.7759/cureus.111193
UR - https://doi.org/10.7759/cureus.111193
LA - en
ER -

CSL-JSON

{
"id": "10.7759/cureus.111193",
"type": "article-journal",
"title": "Benchmarking a Local Schema-Constrained Large Language Model Pipeline for Abstract Screening and Evidence Mapping",
"container-title": "Cureus",
"author": [
{
"family": "Serretti",
"given": "Alessandro"
}
],
"container-title-short": "Cureus",
"volume": "18",
"issue": "6",
"page": "e111193",
"DOI": "10.7759/cureus.111193",
"PMID": "42483140",
"PMCID": "PMC13384908",
"ISSN": "2168-8184",
"publisher": "Cureus Inc.",
"URL": "https://doi.org/10.7759/cureus.111193",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
20
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41398-026-03928-4
Modulatory effects of ketamine on EEG source-based resting state connectivity in treatment resistant depression.
Journal: Translational psychiatry
In common: 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.