OSCR

Health system learning enables generalist neuroimaging models.

Code ↔ Paper

15 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 15 matches
  1. [1] § Methods › NeuroVFM training with Vol-JEPA ↔ neurovfm/systems/pretraining.py, lines 24–55 · score 0.80 · teacher network, student encoder, predictor module, teacher encoder, EMA, JEPA
  2. [2] § Methods › Vision instruction tuning for radiology report generation › NeuroVFM-LLaVA architecture ↔ neurovfm/models/vlm.py, lines 332–474 · score 0.77 · embedding space, connector module, visual tokens, concatenated, Perceiver, scan
  3. [3] § Methods › Diagnostic evaluation of NeuroVFM ↔ neurovfm/systems/classification.py, lines 112–193 · score 0.71 · cross entropy, class weighted, macro, thresholds, Hyperparameters, binary
  4. [4] § Methods › Series preprocessing ↔ neurovfm/data/preprocess.py, lines 167–283 · score 0.69 · Background masks, subdural, clipped, width, axis, blood
  5. [5] § Methods › Evaluation of preliminary report generation ↔ neurovfm/pipelines/interpreter.py, lines 1–22 · score 0.68 · OpenAI, 5–2025, verbosity, medium, GPT, triage
  6. [6] § Learning with Vol-JEPA › NeuroVFM enables preliminary report generation ↔ neurovfm/models/vlm.py, lines 477–496 · score 0.67 · visual instruction tuning, vision language models, LLaVA, style, Multimodal, NeuroImages
  7. [7] § Methods › Grounded diagnoses with multiple instance learning ↔ neurovfm/models/mil.py, lines 390–497 · score 0.66 · attention scores, attention weights, bag, softmax, sum, logits
  8. [8] § Methods › Vision instruction tuning for radiology report generation ↔ neurovfm/models/vlm.py, lines 477–496 · score 0.62 · visual instruction tuning, LLaVA, style, multimodal, architectural, LLM
  9. [9] § Learning with Vol-JEPA › NeuroVFM enables preliminary report generation ↔ neurovfm/systems/llm_sft.py, lines 1–20 · score 0.62 · visual instruction tuning, vision language models, fine tuned, frozen, encoder
  10. [10] § Methods › NeuroVFM training with Vol-JEPA ↔ neurovfm/data/preprocess.py, lines 167–283 · score 0.62 · CT scans, CT window, subdural, crop, axis, bone
  11. [11] § Methods › Grounded diagnoses with multiple instance learning ↔ neurovfm/models/mil.py, lines 282–326 · score 0.60 · standard AB MIL, aggregate, patch, classify, module, models
  12. [12] § Methods › Computational hardware and software ↔ neurovfm/pipelines/diagnostic.py, lines 16–132 · score 0.59 · automatic mixed precision, AMP, CPU, aggregate, Vol, PyTorch
  13. [13] § Methods › Computational hardware and software ↔ neurovfm/pipelines/encoder.py, lines 18–111 · score 0.57 · automatic mixed precision, AMP, CPU, Vol, PyTorch, batch
  14. [14] § Methods › NeuroVFM training with Vol-JEPA ↔ neurovfm/datasets/collators.py, lines 1–19 · score 0.56 · patch dropout, variable length, crops, sequences, batch
  15. [15] § Learning with Vol-JEPA › NeuroVFM enables preliminary report generation ↔ neurovfm/pipelines/interpreter.py, lines 24–102 · score 0.55 · radiologist findings, clinical indication, API, Acuity, NeuroVFM, pipeline

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 857 lines · 33 KB · MIT · 3 matches

  1. """
  2. Vision Language Model for 3D Medical Imaging
  3. Implements a LLaVA-style multimodal model for radiology report generation.
  4. Combines a pretrained vision encoder with a language model via a perceiver-based connector.
  5. Generation follows a JSON schema via outlines/pydantic.
  6. """
  7. import logging
  8. from typing import Any, Dict, List, Optional, Tuple, Union
  9. import outlines
  10. import outlines.caching as cache
  11. import torch
  12. import torch.nn as nn
  13. from outlines.processors.structured import JSONLogitsProcessor
  14. from peft import LoraConfig, TaskType, get_peft_model
  15. from pydantic import BaseModel, Field
  16. from timm.layers.mlp import Mlp
  17. from transformers import (
  18. AutoModelForCausalLM,
  19. AutoTokenizer,
  20. GenerationConfig,
  21. PreTrainedModel,
  22. PreTrainedTokenizer,
  23. )
  24. from neurovfm.models.perceiver import PerceiverResampler
  25. from neurovfm.models.vit import get_vit_backbone
  26. class ShortReport(BaseModel):
  27. # JSON schema for findings generation
  28. exam_type: str = Field(..., description="The type of imaging study.")
  29. findings: List[str] = Field(..., description="A list of key radiological findings.")
  30. class LanguageModel(nn.Module):
  31. """
  32. Wrapper for a pretrained LLM.
  33. """
  34. def __init__(
  35. self,
  36. model_name_or_path: str,
  37. use_gradient_checkpointing: bool = False,
  38. lora_params: Optional[Dict[str, Any]] = None,
  39. attn_implementation: str = "flash_attention_2",
  40. ):
  41. super().__init__()
  42. self.model_name_or_path = model_name_or_path
  43. self.use_gradient_checkpointing = use_gradient_checkpointing
  44. self.lora_params = lora_params
  45. # init LLM and tokenizer
  46. self.llm: PreTrainedModel = AutoModelForCausalLM.from_pretrained(
  47. model_name_or_path,
  48. trust_remote_code=True,
  49. torch_dtype=torch.bfloat16,
  50. low_cpu_mem_usage=True,
  51. attn_implementation=attn_implementation,
  52. )
  53. self.tokenizer = AutoTokenizer.from_pretrained(
  54. model_name_or_path,
  55. trust_remote_code=True,
  56. padding_side='left',
  57. )
  58. self.image_placeholder_token_id = self.tokenizer.convert_tokens_to_ids('<|image_pad|>') # [hardcoded] image padding token used by Qwen2/3 tokenizer
  59. if self.use_gradient_checkpointing:
  60. self.enable_gradient_checkpointing()
  61. if self.lora_params:
  62. self.enable_lora(self.lora_params)
  63. # init structured JSON generator
  64. # monkey-patch tokenizer attributes required for outlines structured generation compatibility
  65. self.tokenizer.vocabulary = self.tokenizer.get_vocab()
  66. self.tokenizer.special_tokens = self.tokenizer.all_special_tokens
  67. self.tokenizer.convert_token_to_string = lambda token: self.tokenizer.decode([self.tokenizer.vocabulary[token]])
  68. cache.disable_cache() # disable caching to prevent OverflowError with large vocabularies
  69. def enable_gradient_checkpointing(self):
  70. """Enable gradient checkpointing for memory efficiency"""
  71. if hasattr(self.llm, 'gradient_checkpointing_enable'):
  72. self.llm.gradient_checkpointing_enable()
  73. def get_input_embeddings(self, input_ids: torch.Tensor) -> torch.Tensor:
  74. """
  75. Get token embeddings from input_ids.
  76. """
  77. embeddings = self.llm.get_input_embeddings()(input_ids)
  78. return embeddings
  79. @property
  80. def hidden_size(self) -> int:
  81. """
  82. Returns the hidden size of the LLM.
  83. """
  84. return self.llm.config.hidden_size
  85. @property
  86. def device(self) -> torch.device:
  87. return self.llm.device
  88. def forward(
  89. self,
  90. input_ids: Optional[torch.Tensor] = None,
  91. attention_mask: Optional[torch.Tensor] = None,
  92. inputs_embeds: Optional[torch.Tensor] = None,
  93. labels: Optional[torch.Tensor] = None,
  94. **kwargs
  95. ) -> Dict[str, torch.Tensor]:
  96. """
  97. Forward pass through the LLM.
  98. """
  99. # ensure consistent dtypes
  100. if inputs_embeds is not None:
  101. inputs_embeds = inputs_embeds.to(torch.bfloat16)
  102. outputs = self.llm(
  103. input_ids=input_ids,
  104. attention_mask=attention_mask,
  105. inputs_embeds=inputs_embeds,
  106. labels=labels,
  107. return_dict=True,
  108. **kwargs
  109. )
  110. return outputs
  111. @torch.inference_mode()
  112. def generate(
  113. self,
  114. inputs_embeds: torch.Tensor,
  115. attention_mask: Optional[torch.Tensor] = None,
  116. **generate_kwargs
  117. ) -> torch.Tensor:
  118. """
  119. Generate text autoregressively.
  120. Args:
  121. inputs_embeds (torch.Tensor): Combined visual and prompt embeddings.
  122. Shape: (batch_size, seq_len, embed_dim)
  123. attention_mask (Optional[torch.Tensor]): Attention mask for inputs_embeds.
  124. Shape: (batch_size, seq_len)
  125. **generate_kwargs: Additional arguments for Hugging Face `generate` method
  126. (e.g., max_length, num_beams, do_sample, temperature).
  127. Returns:
  128. torch.Tensor: Generated token IDs.
  129. """
  130. # ensure consistent dtype
  131. inputs_embeds = inputs_embeds.to(torch.bfloat16)
  132. # pop custom kwarg to avoid passing it to model.generate
  133. schema_str = generate_kwargs.pop('generation_schema', None)
  134. logits_processor = None
  135. if schema_str:
  136. if schema_str == 'shortreport':
  137. schema = ShortReport
  138. else:
  139. raise ValueError(f"Invalid schema: {schema_str}")
  140. json_logits_processor = JSONLogitsProcessor(
  141. schema=schema,
  142. tokenizer=self.tokenizer,
  143. tensor_library_name="torch"
  144. )
  145. logits_processor = [json_logits_processor]
  146. # for structured generation, sampling must be disabled
  147. generate_kwargs['do_sample'] = False
  148. # create a final configuration dictionary by merging model defaults with user kwargs
  149. # User-provided kwargs will override the defaults.
  150. final_config_dict = {
  151. **self.llm.generation_config.to_dict(),
  152. **generate_kwargs
  153. }
  154. # if sampling is disabled (either explicitly or by using num_beams > 1), remove all sampling-related parameters to avoid warnings
  155. if not final_config_dict.get('do_sample', False):
  156. final_config_dict.pop('temperature', None)
  157. final_config_dict.pop('top_p', None)
  158. final_config_dict.pop('top_k', None)
  159. final_config_dict.pop('min_p', None)
  160. final_config_dict.pop('typical_p', None)
  161. # create the final GenerationConfig object from the clean dictionary
  162. generation_config = GenerationConfig.from_dict(final_config_dict)
  163. self.llm.generation_config = generation_config
  164. outputs = self.llm.generate(
  165. inputs_embeds=inputs_embeds,
  166. attention_mask=attention_mask,
  167. logits_processor=logits_processor,
  168. generation_config=generation_config
  169. )
  170. return outputs
  171. def enable_lora(self, lora_config: Optional[Dict[str, Any]] = None):
  172. lora_config['task_type'] = TaskType.CAUSAL_LM
  173. peft_config = LoraConfig(**lora_config)
  174. self.llm = get_peft_model(self.llm, peft_config)
  175. class VisionEncoder(nn.Module):
  176. """
  177. Wrapper for 3D volumetric vision transformer encoder.
  178. Wraps the pretrained VisionTransformer backbone with optional gradient
  179. checkpointing. Outputs per-study or per-batch embeddings.
  180. Note: This module expects pre-normalized inputs. Normalization should be
  181. handled by the pipeline or training system before calling this encoder.
  182. Args:
  183. backbone_cf (Dict): Configuration for get_vit_backbone (which, params)
  184. checkpoint_path (str, optional): Path to pretrained checkpoint
  185. use_gradient_checkpointing (bool): Enable gradient checkpointing for memory efficiency
  186. freeze (bool): Freeze encoder weights. Defaults to True.
  187. Example:
  188. >>> encoder_cf = {
  189. ... 'backbone_cf': {'which': 'vit_base', 'params': {...}},
  190. ... 'checkpoint_path': '/path/to/checkpoint.ckpt',
  191. ... 'freeze': True
  192. ... }
  193. >>> encoder = VisionEncoder(**encoder_cf)
  194. >>> embs, coords = encoder(batch) # batch should have normalized 'img' tokens
  195. """
  196. def __init__(
  197. self,
  198. backbone_cf: Dict[str, Any],
  199. use_gradient_checkpointing: bool = False,
  200. freeze: bool = True,
  201. ):
  202. super().__init__()
  203. # initialize backbone
  204. self.encoder = get_vit_backbone(**backbone_cf)
  205. self.embed_dim = self.encoder.embed_dim
  206. # gradient checkpointing
  207. self.use_gradient_checkpointing = use_gradient_checkpointing
  208. if self.use_gradient_checkpointing:
  209. self._enable_gradient_checkpointing()
  210. # freeze encoder if specified
  211. if freeze:
  212. for param in self.encoder.parameters():
  213. param.requires_grad = False
  214. self.encoder.eval()
  215. def _enable_gradient_checkpointing(self):
  216. """Enable gradient checkpointing for memory efficiency."""
  217. if hasattr(self.encoder, 'set_grad_checkpointing'):
  218. self.encoder.set_grad_checkpointing(True)
  219. elif hasattr(self.encoder, 'blocks'):
  220. for block in self.encoder.blocks:
  221. block.grad_checkpointing = True
  222. def forward(
  223. self,
  224. batch: Dict[str, Any],
  225. return_list: bool = True
  226. ) -> Union[Tuple[List[torch.Tensor], List[torch.Tensor]], Tuple[torch.Tensor, torch.Tensor]]:
  227. """
  228. Forward pass through vision encoder.
  229. Note: Expects pre-normalized 'img' tokens.
  230. Args:
  231. batch (Dict): Batch dictionary with keys:
  232. - img: (total_seqlen, D) normalized image tokens
  233. - coords: (total_seqlen, 3) tensor of 3d coordinates
  234. - series_masks_indices: mask indices for foreground tokens
  235. - series_cu_seqlens: (n_series+1,) cumulative sequence lengths
  236. - series_max_len: int, maximum sequence length
  237. - study_cu_seqlens: (n_studies+1,) study cumulative sequence lengths
  238. return_list (bool): If True, return list of tensors per study.
  239. If False, return single concatenated tensor.
  240. Returns:
  241. If return_list=True:
  242. Tuple[List[Tensor], List[Tensor]]: Per-study embeddings and coordinates
  243. If return_list=False:
  244. Tuple[Tensor, Tensor]: Concatenated embeddings and coordinates
  245. """
  246. # forward pass through encoder
  247. # output shape: (total_tokens_in_batch, embed_dim)
  248. emb = self.encoder.forward(
  249. batch["img"],
  250. batch["coords"],
  251. masks=batch["series_masks_indices"],
  252. cu_seqlens=batch["series_cu_seqlens"],
  253. max_seqlen=batch["series_max_len"]
  254. )
  255. if return_list:
  256. # split embedding/coordinate tensors by study
  257. study_embeddings = []
  258. study_coords_list = []
  259. study_cu_seqlens = batch["study_cu_seqlens"]
  260. for study_idx in range(len(study_cu_seqlens) - 1):
  261. start_idx = study_cu_seqlens[study_idx]
  262. end_idx = study_cu_seqlens[study_idx + 1]
  263. study_emb = emb[start_idx:end_idx]
  264. study_embeddings.append(study_emb.to(torch.bfloat16))
  265. study_coords = batch["coords"][start_idx:end_idx]
  266. study_coords_list.append(study_coords)
  267. return study_embeddings, study_coords_list
  268. else:
  269. return emb.to(torch.bfloat16), batch["coords"]
  270. class VisionConnector(nn.Module):
  271. """
  272. Connector module that projects visual tokens to the LLM embedding space using a perceiver resampler.
  273. Operates on a per-series (scan) basis, compressing each series to a fixed number of tokens before projection.
  274. Args:
  275. visual_embed_dim (int): Dimension of visual encoder embeddings
  276. llm_embed_dim (int): Dimension of LLM embeddings
  277. perceiver_cfg (Dict): Configuration for PerceiverResampler
  278. mlp_hidden_dim (int, optional): Hidden dimension for projection MLP
  279. mlp_drop (float): Dropout rate for MLP. Defaults to 0.0.
  280. mlp_act_layer (nn.Module): Activation layer for MLP. Defaults to nn.GELU.
  281. Example:
  282. >>> connector = VisionConnector(
  283. ... visual_embed_dim=768,
  284. ... llm_embed_dim=4096,
  285. ... perceiver_cfg={'num_queries': 64, 'num_layers': 6, 'num_heads': 8}
  286. ... )
  287. >>> projected, lengths = connector(visual_tokens, series_cu_seqlens)
  288. """
  289. def __init__(
  290. self,
  291. visual_embed_dim: int,
  292. llm_embed_dim: int,
  293. perceiver_cfg: Dict[str, Any],
  294. mlp_hidden_dim: Optional[int] = None,
  295. mlp_drop: float = 0.0,
  296. mlp_act_layer: nn.Module = nn.GELU,
  297. ):
  298. super().__init__()
  299. # build perceiver with visual embed dim
  300. perceiver_cfg = dict(perceiver_cfg)
  301. perceiver_cfg["dim"] = visual_embed_dim
  302. self.perceiver = PerceiverResampler(**perceiver_cfg)
  303. # build projection mlp
  304. self.in_features = visual_embed_dim
  305. self.hidden = mlp_hidden_dim or llm_embed_dim
  306. self.out_features = llm_embed_dim
  307. self.mlp = Mlp(
  308. in_features=self.in_features,
  309. hidden_features=self.hidden,
  310. out_features=self.out_features,
  311. drop=mlp_drop,
  312. act_layer=mlp_act_layer,
  313. )
  314. self.num_queries = self.perceiver.num_queries
  315. self.output_is_list = True
  316. def forward(
  317. self,
  318. visual_tokens: List[torch.Tensor],
  319. serie_cu_seqlens: List[torch.Tensor]
  320. ) -> Tuple[List[torch.Tensor], List[List[int]]]:
  321. """
  322. Compress and project visual tokens.
  323. Args:
  324. visual_tokens (List[Tensor]): List of per-study visual embeddings
  325. Each tensor has shape (N_tokens_for_study, D_vis)
  326. serie_cu_seqlens (List[Tensor]): List of per-study cumulative sequence lengths
  327. Each tensor has shape (N_series + 1,)
  328. Returns:
  329. Tuple[List[Tensor], List[List[int]]]:
  330. - List of projected tokens per study (N_series * num_queries, D_llm)
  331. - List of token counts per series for each study
  332. """
  333. projected_tokens_per_study: List[torch.Tensor] = []
  334. serie_lengths_per_study: List[List[int]] = []
  335. for study_tokens, cu_seqlens in zip(visual_tokens, serie_cu_seqlens):
  336. # study_tokens: (N_total_tokens_for_study, D_vis)
  337. # cu_seqlens: (N_series + 1,) tensor with cumulative lengths
  338. if study_tokens.numel() == 0:
  339. # create placeholder for empty studies (corrupted data)
  340. num_series = len(cu_seqlens) - 1
  341. if num_series > 0:
  342. total_tokens = num_series * self.num_queries
  343. placeholder_tokens = torch.zeros(
  344. total_tokens, self.out_features,
  345. device=self.perceiver.queries.device,
  346. dtype=self.perceiver.queries.dtype
  347. )
  348. placeholder_lengths = [self.num_queries] * num_series
  349. projected_tokens_per_study.append(placeholder_tokens)
  350. serie_lengths_per_study.append(placeholder_lengths)
  351. else:
  352. projected_tokens_per_study.append(
  353. torch.empty(0, self.out_features,
  354. device=self.perceiver.queries.device,
  355. dtype=self.perceiver.queries.dtype)
  356. )
  357. serie_lengths_per_study.append([])
  358. continue
  359. compressed_per_serie = []
  360. num_series = len(cu_seqlens) - 1
  361. # split tokens into series based on cu_seqlens
  362. for i in range(num_series):
  363. start_idx = cu_seqlens[i]
  364. end_idx = cu_seqlens[i + 1]
  365. serie_tokens = study_tokens[start_idx:end_idx]
  366. if serie_tokens.numel() == 0:
  367. # placeholder for empty series
  368. placeholder = torch.zeros(
  369. self.num_queries,
  370. self.in_features,
  371. device=self.perceiver.queries.device,
  372. dtype=self.perceiver.queries.dtype,
  373. )
  374. compressed_per_serie.append(placeholder)
  375. else:
  376. # perceiver expects (B, N, D), so unsqueeze
  377. comp = self.perceiver(serie_tokens.unsqueeze(0)).squeeze(0)
  378. compressed_per_serie.append(comp)
  379. # concatenate compressed tokens from all series
  380. if not compressed_per_serie:
  381. projected_tokens_per_study.append(
  382. torch.empty(0, self.out_features,
  383. device=study_tokens.device,
  384. dtype=study_tokens.dtype)
  385. )
  386. serie_lengths_per_study.append([])
  387. continue
  388. all_serie_compressed = torch.cat(compressed_per_serie, dim=0)
  389. # project to llm dimension
  390. projected = self.mlp(all_serie_compressed)
  391. projected_tokens_per_study.append(projected)
  392. serie_lengths_per_study.append([self.num_queries] * num_series)
  393. return projected_tokens_per_study, serie_lengths_per_study
  394. class VisionLanguageModel(nn.Module):
  395. """
  396. Multimodal LLM for neuroimaging report generation.
  397. Combines a pretrained vision encoder, perceiver-based connector, and a language model in a LLaVA-style architecture for visual instruction tuning.
  398. Args:
  399. vision_encoder_cf (Dict): Configuration for VisionEncoder
  400. vision_connector_cf (Dict): Configuration for VisionConnector
  401. language_model_cf (Dict): Configuration for LanguageModel
  402. use_gradient_checkpointing (bool): Enable gradient checkpointing
  403. Example:
  404. >>> model = NeuroLlavaModel(
  405. ... vision_encoder_cf={...},
  406. ... vision_connector_cf={...},
  407. ... language_model_cf={...}
  408. ... )
  409. >>> outputs = model(vision_batch, input_ids, attention_mask, labels)
  410. """
  411. def __init__(
  412. self,
  413. vision_encoder_cf: Dict[str, Any],
  414. vision_connector_cf: Dict[str, Any],
  415. language_model_cf: Dict[str, Any],
  416. use_gradient_checkpointing: bool = False,
  417. ):
  418. super().__init__()
  419. # initialize vision encoder
  420. self.vision_encoder = VisionEncoder(
  421. backbone_cf=vision_encoder_cf,
  422. use_gradient_checkpointing=use_gradient_checkpointing,
  423. freeze=True,
  424. ).to(torch.bfloat16)
  425. self.visual_embed_dim = self.vision_encoder.embed_dim
  426. # initialize language model
  427. attn_impl = language_model_cf.get("attn_implementation", "flash_attention_2")
  428. self.language_model = LanguageModel(
  429. model_name_or_path=language_model_cf['model_name_or_path'],
  430. use_gradient_checkpointing=use_gradient_checkpointing,
  431. lora_params=language_model_cf.get('lora_params', None), # optional LoRA parameters; if None, finetunes entire LLM
  432. attn_implementation=attn_impl,
  433. ).to(torch.bfloat16)
  434. llm_dim = self.language_model.hidden_size
  435. # initialize vision connector
  436. cfg = vision_connector_cf.get('serie_perceiver', vision_connector_cf)
  437. self.vision_connector = VisionConnector(
  438. visual_embed_dim=self.visual_embed_dim,
  439. llm_embed_dim=llm_dim,
  440. perceiver_cfg=cfg.get('perceiver_cfg', cfg),
  441. mlp_hidden_dim=vision_connector_cf.get('mlp_hidden_dim'),
  442. mlp_drop=vision_connector_cf.get('mlp_drop', 0.0),
  443. ).to(torch.bfloat16)
  444. def forward(
  445. self,
  446. vision_batch: Optional[Dict[str, Any]],
  447. input_ids: Optional[torch.Tensor],
  448. attention_mask: Optional[torch.Tensor],
  449. labels: Optional[torch.Tensor] = None,
  450. **kwargs,
  451. ) -> Union[Dict[str, torch.Tensor], torch.Tensor]:
  452. """
  453. Forward pass for training or generation.
  454. Args:
  455. vision_batch (Dict): Batch from StudyPreprocessor
  456. input_ids (Tensor): Tokenized input ids (B, S)
  457. attention_mask (Tensor): Attention mask (B, S)
  458. labels (Tensor, optional): Labels for training (B, S)
  459. Returns:
  460. CausalLMOutputWithPast for training, or generated ids for inference
  461. """
  462. # 1. get projected visual embeddings
  463. vision_output = self._forward_vision(vision_batch)
  464. # 2. get text embeddings
  465. text_embeds = self.language_model.get_input_embeddings(input_ids)
  466. # 3. splice visual and text embeddings
  467. inputs_embeds, attention_mask, labels = self._splice_vision_and_text(
  468. vision_output=vision_output,
  469. text_embeds=text_embeds,
  470. input_ids=input_ids,
  471. attention_mask=attention_mask,
  472. labels=labels,
  473. )
  474. # 4. forward through language model
  475. return self._forward_sft(
  476. inputs_embeds=inputs_embeds,
  477. attention_mask=attention_mask,
  478. labels=labels,
  479. )
  480. def generate(
  481. self,
  482. vision_batch: Optional[Dict[str, Any]],
  483. input_ids: Optional[torch.Tensor],
  484. attention_mask: Optional[torch.Tensor],
  485. **generate_kwargs,
  486. ):
  487. """
  488. Autoregressive text generation.
  489. Args:
  490. vision_batch (Dict): Batch from StudyPreprocessor
  491. input_ids (Tensor): Tokenized prompt ids
  492. attention_mask (Tensor): Attention mask for prompt
  493. **generate_kwargs: HuggingFace generate arguments
  494. Returns:
  495. Generated token ids or dict with hidden states if requested
  496. """
  497. # 1. get projected visual embeddings
  498. vision_output = self._forward_vision(vision_batch)
  499. # 2. get text embeddings
  500. text_embeds = self.language_model.get_input_embeddings(input_ids)
  501. # 3. splice visual and text embeddings
  502. inputs_embeds, attention_mask, _ = self._splice_vision_and_text(
  503. vision_output=vision_output,
  504. text_embeds=text_embeds,
  505. input_ids=input_ids,
  506. attention_mask=attention_mask,
  507. labels=None,
  508. )
  509. # 4. generate
  510. return self._forward_generate(
  511. inputs_embeds=inputs_embeds,
  512. attention_mask=attention_mask,
  513. **generate_kwargs,
  514. )
  515. def _splice_vision_and_text(
  516. self,
  517. vision_output,
  518. text_embeds,
  519. input_ids,
  520. attention_mask,
  521. labels=None,
  522. ) -> Tuple[torch.Tensor, torch.Tensor, Optional[torch.Tensor]]:
  523. """
  524. Splice visual embeddings into text at placeholder locations.
  525. Args:
  526. vision_output: Output from _forward_vision
  527. text_embeds (Tensor): Text embeddings (B, S_text, D)
  528. input_ids (Tensor): Input ids (B, S_text)
  529. attention_mask (Tensor): Attention mask (B, S_text)
  530. labels (Tensor, optional): Labels (B, S_text)
  531. Returns:
  532. Tuple of (inputs_embeds, attention_mask, labels), all left-padded
  533. """
  534. final_embeds, final_attention_mask = [], []
  535. final_labels = [] if labels is not None else None
  536. # handle multiple placeholders for per-serie vision connectors
  537. projected_visual_tokens, projected_visual_serie_lengths = vision_output
  538. for i in range(text_embeds.shape[0]):
  539. study_visual_tokens = projected_visual_tokens[i]
  540. serie_lengths = projected_visual_serie_lengths[i]
  541. visual_token_chunks = list(torch.split(study_visual_tokens, serie_lengths, dim=0))
  542. placeholder_indices = (input_ids[i] == self.language_model.image_placeholder_token_id).nonzero(as_tuple=True)[0]
  543. if len(placeholder_indices) != len(visual_token_chunks):
  544. raise ValueError(
  545. f"Mismatch between number of image placeholders and image chunks for sample {i}.",
  546. f"Number of image placeholders: {len(placeholder_indices)}",
  547. f"Number of image chunks: {len(visual_token_chunks)}",
  548. )
  549. spliced_embeds_parts, spliced_mask_parts = [], []
  550. spliced_labels_parts = [] if labels is not None else None
  551. last_text_idx = 0
  552. for j, placeholder_idx in enumerate(placeholder_indices):
  553. spliced_embeds_parts.append(text_embeds[i, last_text_idx:placeholder_idx])
  554. spliced_mask_parts.append(attention_mask[i, last_text_idx:placeholder_idx])
  555. if labels is not None:
  556. spliced_labels_parts.append(labels[i, last_text_idx:placeholder_idx])
  557. visual_chunk = visual_token_chunks[j]
  558. spliced_embeds_parts.append(visual_chunk)
  559. spliced_mask_parts.append(torch.ones(visual_chunk.shape[0], dtype=torch.long, device=text_embeds.device))
  560. if labels is not None:
  561. spliced_labels_parts.append(torch.full((visual_chunk.shape[0],), -100, dtype=torch.long, device=text_embeds.device))
  562. last_text_idx = placeholder_idx + 1
  563. spliced_embeds_parts.append(text_embeds[i, last_text_idx:])
  564. spliced_mask_parts.append(attention_mask[i, last_text_idx:])
  565. if labels is not None:
  566. spliced_labels_parts.append(labels[i, last_text_idx:])
  567. final_embeds.append(torch.cat(spliced_embeds_parts, dim=0))
  568. final_attention_mask.append(torch.cat(spliced_mask_parts, dim=0))
  569. if final_labels is not None:
  570. final_labels.append(torch.cat(spliced_labels_parts, dim=0))
  571. # left pad all sequences to the same length
  572. batch_size = len(final_embeds)
  573. max_len = max(len(s) for s in final_embeds)
  574. embed_dim = final_embeds[0].shape[-1]
  575. device, dtype = final_embeds[0].device, final_embeds[0].dtype
  576. inputs_embeds = torch.zeros(batch_size, max_len, embed_dim, device=device, dtype=dtype)
  577. attention_mask_padded = torch.zeros(batch_size, max_len, dtype=torch.long, device=device)
  578. for i, seq in enumerate(final_embeds):
  579. seq_len = len(seq)
  580. inputs_embeds[i, -seq_len:] = seq
  581. attention_mask_padded[i, -seq_len:] = final_attention_mask[i]
  582. labels_padded = None
  583. if final_labels is not None:
  584. labels_padded = torch.full((batch_size, max_len), -100, dtype=torch.long, device=device)
  585. for i, seq in enumerate(final_labels):
  586. labels_padded[i, -len(seq):] = seq
  587. return inputs_embeds, attention_mask_padded, labels_padded
  588. def _forward_sft(
  589. self,
  590. inputs_embeds,
  591. attention_mask,
  592. labels,
  593. ):
  594. """Forward pass for supervised fine-tuning."""
  595. if labels is None:
  596. raise ValueError("For training, 'labels' must be provided.")
  597. return self.language_model(
  598. inputs_embeds=inputs_embeds,
  599. attention_mask=attention_mask,
  600. labels=labels
  601. )
  602. def _forward_generate(
  603. self,
  604. inputs_embeds,
  605. attention_mask,
  606. **generate_kwargs
  607. ):
  608. """Forward pass for generation."""
  609. return self.language_model.generate(
  610. inputs_embeds=inputs_embeds,
  611. attention_mask=attention_mask,
  612. **generate_kwargs
  613. )
  614. def _forward_vision(
  615. self,
  616. vision_batch: Optional[Dict[str, Any]],
  617. ) -> Union[torch.Tensor, List[torch.Tensor], Tuple[List[torch.Tensor], List[List[int]]]]:
  618. """
  619. Forward pass through vision encoder and connector.
  620. Args:
  621. vision_batch (Dict): Batch from StudyPreprocessor
  622. Returns:
  623. Projected visual tokens (format depends on connector strategy)
  624. """
  625. visual_tokens, visual_coords = self.vision_encoder(vision_batch)
  626. # need to create a series_cu_seqlens_list, where each tensor in the list is the cumulative lengths of the series for a given study
  627. # this is used by the vision connector to apply perceiver in a series-wise manner
  628. # example:
  629. # study_cu_seqlens = [0, 1000, 2000, 3000]
  630. # series_cu_seqlens = [0, 250, 1000, 1250, 1500, 2000, 2500, 3000]
  631. # series_cu_seqlens_list = [[0, 250, 1000], [0, 250, 500, 1000], [0, 500, 1000]]
  632. study_cu_seqlens = vision_batch['study_cu_seqlens']
  633. series_cu_seqlens_flat = vision_batch['series_cu_seqlens']
  634. batch_size = len(study_cu_seqlens) - 1
  635. series_cu_seqlens_list = []
  636. start_idx_in_series_flat = 0
  637. for i in range(batch_size):
  638. study_start_offset = study_cu_seqlens[i]
  639. study_end_offset = study_cu_seqlens[i+1]
  640. # find the index of the study's end offset in the flat series tensor
  641. end_idx_tensor = (series_cu_seqlens_flat == study_end_offset).nonzero(as_tuple=True)[0]
  642. if end_idx_tensor.numel() == 0:
  643. raise ValueError(f"Study boundary {study_end_offset} not found in series_cu_seqlens.")
  644. if len(end_idx_tensor) > 1:
  645. raise ValueError(f"Multiple indices found for study boundary {study_end_offset} in series_cu_seqlens: {end_idx_tensor}\nstudy_cu_seqlens: {study_cu_seqlens}\nseries_cu_seqlens_flat: {series_cu_seqlens_flat}")
  646. end_idx_in_series_flat = end_idx_tensor.item()
  647. # slice the cumulative lengths for the current study
  648. study_series_cu_seqlens = series_cu_seqlens_flat[start_idx_in_series_flat : end_idx_in_series_flat + 1]
  649. # make the lengths relative to the start of the study
  650. study_series_cu_seqlens_relative = study_series_cu_seqlens - study_start_offset
  651. series_cu_seqlens_list.append(study_series_cu_seqlens_relative)
  652. # the start for the next study is the end of the current one
  653. start_idx_in_series_flat = end_idx_in_series_flat
  654. return self.vision_connector(visual_tokens, series_cu_seqlens_list)
  655. def _prepare_inputs_for_generation(
  656. self,
  657. input_ids,
  658. attention_mask,
  659. labels,
  660. ) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor, List[str]]:
  661. """
  662. Extract prompt from tokenized conversation for generation.
  663. Args:
  664. input_ids (Tensor): Full tokenized conversation (B, S)
  665. attention_mask (Tensor): Attention mask (B, S)
  666. labels (Tensor): Labels with -100 for prompt tokens (B, S)
  667. Returns:
  668. Tuple of (input_ids, attention_mask, labels, decoded_texts)
  669. """
  670. # extract prompt parts from each sample in the batch
  671. # since data is left-padded, we can directly slice from start to first non -100 label
  672. gen_input_ids_list = []
  673. max_prompt_len = 0
  674. device = input_ids.device
  675. for i in range(len(input_ids)):
  676. # find the start of the response (first non -100 label)
  677. labels_i = labels[i]
  678. first_target_idx_tensor = (labels_i != -100).nonzero(as_tuple=True)[0]
  679. if len(first_target_idx_tensor) > 0:
  680. first_target_idx = first_target_idx_tensor[0]
  681. # the prompt is everything up to the start of the response
  682. prompt_ids = input_ids[i][:first_target_idx]
  683. else:
  684. # if no target, the whole sequence is the prompt (can happen with truncation)
  685. prompt_ids = input_ids[i]
  686. gen_input_ids_list.append(prompt_ids)
  687. if len(prompt_ids) > max_prompt_len:
  688. max_prompt_len = len(prompt_ids)
  689. # left pad all input_ids to max_prompt_len
  690. padded_gen_input_ids_list = []
  691. for prompt_ids in gen_input_ids_list:
  692. pad_len = max_prompt_len - len(prompt_ids)
  693. padded = torch.cat([
  694. torch.full((pad_len,), self.language_model.tokenizer.pad_token_id, dtype=torch.long, device=device),
  695. prompt_ids
  696. ])
  697. padded_gen_input_ids_list.append(padded)
  698. # decode back to raw text (for debugging)
  699. generation_texts = [self.language_model.tokenizer.decode(padded, skip_special_tokens=True).strip() for padded in padded_gen_input_ids_list]
  700. # return the original batch with the updated values
  701. input_ids_new = torch.stack(padded_gen_input_ids_list)
  702. attention_mask_new = torch.ones_like(input_ids_new).to(device)
  703. labels_new = None
  704. return input_ids_new, attention_mask_new, labels_new, generation_texts

vlm.py at commit eb41760, under MIT · at the source

Overview

Authors: Akhil Kondepudi1,2, Akshay Rao1, Chenhui Zhao1,3, Yiwei Lyu1,3, Samir Harake1, Soumyanil Banerjee1, Jacob Ogle1, Rushikesh Joshi1,4, Anna-Katharina Meissner5, Xinhai Hou1,2, Cheng Jiang1,2, Asadur Chowdury1, Ashok Srinivasan6, Brian Athey2, Vikas Gulani6, Aditya Pandey4, Honglak Lee3, Todd Hollon1,2,3,4
  1. Machine Learning in Neurosurgery Lab, University of Michigan, Ann Arbor, MI USA
  2. University of Michigan Computational Medicine and Bioinformatics, Ann Arbor, MI USA
  3. University of Michigan Computer Science and Engineering, Ann Arbor, MI USA
  4. University of Michigan Neurosurgery, Ann Arbor, MI USA
  5. University of Cologne Neurosurgery, Cologne, Germany
  6. University of Michigan Radiology, Ann Arbor, MI USA
Institutions: University of Michigan (United States); University of Cologne (Germany); University Hospital Cologne (Germany); Michigan Medicine (United States)
Journal: Nature medicine, volume 32, issue 8, pages 2831-2837
Dates: received 12 March 2026; accepted 29 May 2026; published online 10 July 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41591-026-04497-1 · PMID 42432292 · PMCID PMC13472962 · OpenAlex W7167936829
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality), human (organism), clinical / translational (subfield)
Methods: Connectivity, Statistics, Machine learning, Smoothing, state filtering, decompositions
Keywords: Machine learning, Translational research, Image processing, Biomedical engineering
MeSH: Artificial Intelligence*, Learning Health System*, Neuroimaging*, Brain, Humans, Magnetic Resonance Imaging, Tomography, X-Ray Computed (* major topic)
Topic: Artificial Intelligence in Healthcare and Education (Health Informatics, Medicine), according to OpenAlex
Funding: Ian&apos;s Friends Foundation; University of Michigan, Chan Zuckerberg Foundation; U.S. Department of Health &amp; Human Services | NIH | National Institute of Neurological Disorders and Stroke; U.S. Department of Health &amp; Human Services | NIH | National Cancer Institute
Citations: cited by 1 paper (Europe PMC); 59 references in the paper

Abstract

Frontier artificial intelligence (AI) models have advanced rapidly through training on internet-scale public data, yet such systems lack access to private clinical data. Neuroimaging is underrepresented in the public domain due to identifiable facial features within magnetic resonance imaging (MRI) and computed tomography (CT) scans, restricting model performance in clinical medicine. Here we show that frontier models underperform on neuroimaging tasks and that learning directly from uncurated data generated during routine clinical care at health systems, a paradigm we call ‘health system learning’, yields high-performance, generalist neuroimaging models. We introduce NeuroVFM, a visual foundation model trained on 5.24 million clinical MRI and CT volumes using a scalable volumetric predictive architecture. NeuroVFM learns comprehensive representations of brain anatomy and pathology, achieving state-of-the-art performance across multiple clinical tasks, including radiologic diagnosis and report generation. The model embeds MRI and CT scans into a shared neuroanatomic latent space and grounds diagnostic findings. When paired with open-source language models, NeuroVFM generates radiology reports that surpass frontier models in accuracy, clinical triage and expert preference. NeuroVFM reduces hallucinated findings and critical errors, offering safer clinical decision support. These results establish health system learning as a paradigm for building generalist medical AI and provide a scalable framework for clinical foundation models.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.

MLNeurosurg/neurovfm

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: eb41760f200e3a542fd7bce627513d9a625f0997, 10 September 2026
Languages: Python (55)
Size: 73 files, 55 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (pyproject.toml), tests
Not found: CITATION.cff, continuous integration, documentation
Tools: PyTorch (40 files), NumPy (10 files), PyTorch Lightning (6 files), SimpleITK (3 files), pandas (2 files), PyTorch Geometric (2 files), Hugging Face Transformers (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
57 files

Code availability

All code was implemented in Python (version 3.10.14) using PyTorch (2.5.0) compiled with CUDA 12.4 as the primary machine learning framework. The following packages were used for data preprocessing, model training and evaluation: pydicom (2.4.4), nibabel (5.3.2), SimpleITK (2.4.0), torchvision (0.20.0), pandas (2.2.3), NumPy (2.1.2), PyTorch Lightning (2.5.0.post0), flash-attn (2.6.3), matplotlib (3.10.7), scipy (1.15.2) and scikit-learn (1.6.1). The following packages were used to load baselines: open-clip (2.23.0) and transformers (4.56.0). All code and scripts to reproduce the experiments in this study are available on GitHub at https://github.com/MLNeurosurg/neurovfm under an MIT license.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 55 scripts, each with its path and the digest of its content;
  • 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

IRB approval was obtained from the University of Michigan for MRI data collection. Restrictions apply to the availability of raw patient MRI and CT imaging data, which were used with institutional permission through IRB approval for the present study and, thus, are not publicly available. All data sharing among medical centers is regulated through data use agreements with the study authors. A similar data-sharing protocol may be established for interested investigators. Please contact the corresponding author (T.H.) for any requests for data sharing. All requests will be evaluated based on institutional and departmental policies to determine whether the data requested are subject to intellectual property or patient privacy obligations. Data can be shared only for non-commercial academic and investigational purposes.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 18 authors, 4 keywords, 7 MeSH terms, 4 funders, 32 references.

Cite

This paper

Kondepudi, A., Rao, A., Zhao, C., Lyu, Y., Harake, S., Banerjee, S., Ogle, J., Joshi, R., Meissner, A.-K., Hou, X., Jiang, C., Chowdury, A., Srinivasan, A., Athey, B., Gulani, V., Pandey, A., Lee, H., & Hollon, T. (2026). Health system learning enables generalist neuroimaging models. Nature medicine, 32(8), 2831-2837. https://doi.org/10.1038/s41591-026-04497-1

BibTeX

@article{kondepudi2026health,
author = {Kondepudi, Akhil and Rao, Akshay and Zhao, Chenhui and Lyu, Yiwei and Harake, Samir and Banerjee, Soumyanil and Ogle, Jacob and Joshi, Rushikesh and Meissner, Anna-Katharina and Hou, Xinhai and Jiang, Cheng and Chowdury, Asadur and Srinivasan, Ashok and Athey, Brian and Gulani, Vikas and Pandey, Aditya and Lee, Honglak and Hollon, Todd},
title = {{Health system learning enables generalist neuroimaging models}},
journal = {Nature medicine},
year = {2026},
month = jul,
volume = {32},
number = {8},
pages = {2831--2837},
publisher = {Nature Portfolio},
issn = {1078-8956},
doi = {10.1038/s41591-026-04497-1},
url = {https://doi.org/10.1038/s41591-026-04497-1},
pmid = {42432292},
pmcid = {PMC13472962}
}

RIS

TY - JOUR
AU - Kondepudi, Akhil
AU - Rao, Akshay
AU - Zhao, Chenhui
AU - Lyu, Yiwei
AU - Harake, Samir
AU - Banerjee, Soumyanil
AU - Ogle, Jacob
AU - Joshi, Rushikesh
AU - Meissner, Anna-Katharina
AU - Hou, Xinhai
AU - Jiang, Cheng
AU - Chowdury, Asadur
AU - Srinivasan, Ashok
AU - Athey, Brian
AU - Gulani, Vikas
AU - Pandey, Aditya
AU - Lee, Honglak
AU - Hollon, Todd
TI - Health system learning enables generalist neuroimaging models
T2 - Nature medicine
J2 - Nat Med
PY - 2026
DA - 2026/07/10
VL - 32
IS - 8
SP - 2831
EP - 2837
SN - 1078-8956
PB - Nature Portfolio
DO - 10.1038/s41591-026-04497-1
UR - https://doi.org/10.1038/s41591-026-04497-1
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41591-026-04497-1",
"type": "article-journal",
"title": "Health system learning enables generalist neuroimaging models",
"container-title": "Nature medicine",
"author": [
{
"family": "Kondepudi",
"given": "Akhil"
},
{
"family": "Rao",
"given": "Akshay"
},
{
"family": "Zhao",
"given": "Chenhui"
},
{
"family": "Lyu",
"given": "Yiwei"
},
{
"family": "Harake",
"given": "Samir"
},
{
"family": "Banerjee",
"given": "Soumyanil"
},
{
"family": "Ogle",
"given": "Jacob"
},
{
"family": "Joshi",
"given": "Rushikesh"
},
{
"family": "Meissner",
"given": "Anna-Katharina"
},
{
"family": "Hou",
"given": "Xinhai"
},
{
"family": "Jiang",
"given": "Cheng"
},
{
"family": "Chowdury",
"given": "Asadur"
},
{
"family": "Srinivasan",
"given": "Ashok"
},
{
"family": "Athey",
"given": "Brian"
},
{
"family": "Gulani",
"given": "Vikas"
},
{
"family": "Pandey",
"given": "Aditya"
},
{
"family": "Lee",
"given": "Honglak"
},
{
"family": "Hollon",
"given": "Todd"
}
],
"container-title-short": "Nat Med",
"volume": "32",
"issue": "8",
"page": "2831-2837",
"DOI": "10.1038/s41591-026-04497-1",
"PMID": "42432292",
"PMCID": "PMC13472962",
"ISSN": "1078-8956",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41591-026-04497-1",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1162/imag.a.1326 [code]
RAVEN: Robust, generalizable, multi-resolution structural MRI upsampling using autoencoders.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: PyTorch Lightning, Hugging Face Transformers, SimpleITK, 3 other tools, structural MRI / diffusion, 1 reference
[2] doi:10.1093/radadv/umag025 [code]
OpenMAP-BrainAge: generalizable and interpretable brain age predictor from MRI.
Journal: Radiology advances
In common: Hugging Face Transformers, SimpleITK, PyTorch, 2 other tools, structural MRI / diffusion, 1 reference
[3] doi:10.1016/j.patter.2026.101538 [code]
A multi-modal foundation model for brain disease diagnosis and medical imaging.
Journal: Patterns (New York, N.Y.)
In common: Hugging Face Transformers, PyTorch, pandas, 1 other tool, clinical / translational, 2 references
[4] doi:10.1038/s44387-026-00109-y [code]
SPARROW: subtyping Parkinson's disease with agentic reasoning and robust omics workflow.
Journal: NPJ artificial intelligence
In common: Hugging Face Transformers, PyTorch, pandas, 1 other tool, clinical / translational, 2 references
[5] doi:10.1002/hbm.70469 [code]
VarCoNet: A Variability-Aware Self-Supervised Framework for Functional Connectome Extraction From Resting-State fMRI.
Journal: Human brain mapping
In common: PyTorch Geometric, Hugging Face Transformers, PyTorch, 2 other tools, 1 reference
[6] doi:10.1038/s43856-026-01722-3 [code]
Local and global patterns support medical imaging as a biomarker of ageing.
Journal: Communications medicine
In common: PyTorch Lightning, SimpleITK, PyTorch, 2 other tools, clinical / translational
[7] doi:10.1016/j.isci.2026.117446 [code]
Volume-inflation registration (INFREG) for morphometric analysis of human focal cortical dysplasia type II.
Journal: iScience
In common: PyTorch Geometric, SimpleITK, PyTorch, 2 other tools, structural MRI / diffusion
[8] doi:10.1186/s41747-026-00735-w [code]
Dual-conditioned diffusion model with anatomical guidance for geometric distortion correction in prostate MRI.
Journal: European radiology experimental
In common: Hugging Face Transformers, SimpleITK, PyTorch, 2 other tools, structural MRI / diffusion
[9] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: PyTorch Lightning, PyTorch, pandas, 1 other tool, structural MRI / diffusion, clinical / translational, 1 reference
[10] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: PyTorch Lightning, PyTorch Geometric, PyTorch, 2 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.