The Brain Imaging and Neurophysiology Dataset of large-scale multimodal neural data.
The 4 matches
- [1] § Methods › Computational processing › Sequence names ↔ mri_sequence_identification.ipynb, lines 95–223 · score 0.99 · Echo Train Length, Flip Angle, fMRI, MR Acquisition, Sequence Variant, MRI sequence
- [2] § Methods › Computational processing › Clinical metadata extraction ↔ ask_llama_open.py, lines 153–233 · score 0.71 · mentioned pathologies, brain related, clinically important, acute, chronic, location
- [3] § Methods › Computational processing › Clinical metadata extraction ↔ ask_llama_category.py, lines 72–121 · score 0.64 · Bio Medical Llama, 3–8, asked, answer, Model, Clinical
- [4] § Methods › Computational processing › Clinical metadata extraction ↔ ask_llama_open.py, lines 102–151 · score 0.60 · Bio Medical Llama, 3–8, asked, Model
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 233 lines · 9.3 KB · CC-BY-NC-4.0 · 2 matches
- import os
- import json
- import torch
- import tarfile
- import argparse
- import numpy as np
- import pandas as pd
- from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
- def clean_json_string(json_str):
- # Try to load the string into a Python object
- try:
- # Attempt to load and parse the string, if valid JSON, return it
- data = json.loads(json_str)
- return data
- except json.JSONDecodeError:
- try:
- # If an error occurs, the string is likely incomplete, Remove the part after the last complete JSON object or array
- last_complete_json_end = json_str.rfind('}')
- if last_complete_json_end != -1:
- fixed_str = json_str[:last_complete_json_end + 1]
- return json.loads(fixed_str)
- except:
- # erroreous output
- new_str = '{"Name": "INCOMPLETE"}'
- return json.loads(new_str)
- def ask_llama (system_message,instruction_message, pipeline):
- messages = [{"role": "system", "content": system_message},
- {"role": "user",
- "content": instruction_message},
- ]
- prompt = pipeline.tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
- terminators = [pipeline.tokenizer.eos_token_id, pipeline.tokenizer.convert_tokens_to_ids("<|eot_id|>")]
- with torch.no_grad():
- outputs = pipeline(prompt, max_new_tokens=2000, eos_token_id=terminators, do_sample=True, temperature=0.5,
- top_p=0.9)
- return outputs
- def read_from_tar(tar_path, file_name):
- with tarfile.open(tar_path, 'r') as tar:
- # Check if the file exists in the tar archive
- if file_name in tar.getnames():
- file = tar.extractfile(file_name)
- file_content = file.read()
- file_content = file_content.decode('utf-8')
- file_content = eval(file_content) # to dict
- else:
- print(f"File '{file_name}' not found in the tar archive.")
- file_content = None
- return file_content
- if __name__ == "__main__":
- parser = argparse.ArgumentParser(description="Run distance analysis")
- parser.add_argument(
- "-tar_data_dir",
- type=str,
- action="store",
- help="path to tar directory with data",
- )
- parser.add_argument(
- "-master_dir",
- action="store",
- default=False,
- help="Directory to csv with all patients",
- )
- parser.add_argument(
- "-output_dir",
- action="store",
- default=False,
- help="Directory to store results in",
- )
- parser.add_argument(
- "-model_dir",
- action="store",
- default=False,
- help="Directory to with models",
- )
- parser.add_argument(
- "-reverse",
- action="store",
- default=False,
- help="0 or 1 Start at the End of the Dataset to better parallelize",
- )
- parser.add_argument(
- "-split",
- action="store",
- default=False,
- help="between 0 to 80000, start index to calculate ten thousand samples.",
- )
- # set up device
- device = "cpu"
- if torch.cuda.is_available():
- device = "cuda"
- # if torch.backends.mps.is_available():
- # device = "mps"
- print(f'USING DEVICE {device}')
- # load data
- args = parser.parse_args()
- master = args.master_dir
- reverse = args.reverse
- REPORT_DIR = args.tar_data_dir
- OUT_DIR = args.output_dir
- model_dir = args.model_dir
- split = int(args.split)
- os.makedirs(OUT_DIR, exist_ok=True)
- master = pd.read_csv(master)
- # load llama
- model_name = 'ContactDoctor/Bio-Medical-Llama-3-8B'
- tokenizer = AutoTokenizer.from_pretrained(f'{model_dir}/tokenizers/{model_name}')
- model = AutoModelForCausalLM.from_pretrained(f'{model_dir}/models/{model_name}',device_map=device)
- pipel = pipeline("text-generation", model=model,tokenizer=tokenizer, model_kwargs={"torch_dtype": torch.bfloat16})
- # set model to eval mode
- model.eval()
- # if MGB
- if 'MGB' in REPORT_DIR:
- master['aws_file'] = master['aws_filename'].str.replace('/','_')
- if split:
- selection = master['aws_file'][split:split+10000]
- else:
- selection = master['aws_file']
- # start at the end if reverese = 1
- if int(reverse):
- selection = selection.iloc[::-1]
- for id_stan in selection:
- id_tmp = id_stan.split('.')[0]
- # fname = f"{REPORT_DIR}/Report_ID_{id_stan[:-4]}.json"
- out_file1 = f'{OUT_DIR}/LLAMA_out_{id_tmp}_scaninfo.json'
- out_file2 = f'{OUT_DIR}/LLAMA_out_{id_tmp}_findings.json'
- folder_name = REPORT_DIR.split('/')[-1][:-7]
- report = f'{folder_name}/Report_ID_{id_tmp}.json'
- # if one of the output does not exist yet:
- if not os.path.isfile(out_file1) or not os.path.isfile(out_file2):
- data_id = read_from_tar(REPORT_DIR, report)
- # if data_id not found in data
- if not data_id:
- data_id = 0 # do not process this one.
- # ID NOT FOUND
- with open(out_file1, "w") as out:
- json.dump("REPORT NOT FOUND", out)
- with open(out_file2, "w") as out:
- json.dump("REPORT NOT FOUND", out)
- # if data_id found in tar
- elif data_id:
- answers = dict()
- # set input for model
- question = dict()
- question['Brain']="Is this report about the brain? Answer 'Yes', 'No' or 'Unknown'."
- question['Abnormality']="Is this a pathological report? Answer 'Yes', 'No' or 'Unknown'."
- question['ALL-free']= "Make a list of all mentioned pathologies in the report, start with the most clinically important one. Do not repeat pathologies. Only include pathologies which are explicitly mentioned in the report. Return a list of pathologies in JSON format. Each finding should have a 'Name', 'General Name', 'Location', 'Brain-related', 'Magnitude','Acute/Chronic' and a 'Details' field."
- system_message = f"You are an expert trained on neurology and clinical domain. Do only report information which is explicitely metnioed in the report."
- report = data_id['report']
- answers['report'] = report
- for qu in ['Brain','Abnormality']:
- instruction_message = f'{report} {question[qu]}'
- answer_tmp = ask_llama(system_message,instruction_message, pipel)
- answer_short = answer_tmp[0]['generated_text'].split("<|eot_id|>\n\nAssistant: ")[-1]
- answers[qu] = answer_short
- print(f'Finished question {qu}. Response: {answer_short}')
- # save participant_dict with scan and brain info
- with open(out_file1, "w") as out:
- json.dump(answers, out)
- # now ask for the actual findings in the report
- for qu in ['ALL-free']:
- instruction_message = f'{report} {question[qu]}'
- answer_tmp = ask_llama(system_message,instruction_message, pipel)
- answer_short = answer_tmp[0]['generated_text'].split("<|eot_id|>\n\nAssistant: ")[-1]
- findings = answer_short
- print(f'Finished question {qu}. Response: {findings}')
- findings_list = clean_json_string(findings)
- new_findings = []
- if isinstance(findings_list, list):
- if len(findings_list)>=1:
- # Second round of LLAMA to self-correct
- for f_tmp in findings_list:
- try:
- # ask it to self-correct
- instruction_message = f"{report} Is {f_tmp['General Name']} explicitly mentioned in the report? Answer Yes or No"
- answer_tmp = ask_llama(system_message,instruction_message, pipel)
- answer_short = answer_tmp[0]['generated_text'].split("<|eot_id|>\n\nAssistant: ")[-1]
- f_tmp['Test'] = answer_short
- except:
- f_tmp['Test'] = 'UNKNOWN'
- try:
- # remove names and dates from the term
- instruction_message = f"Does the following text contain names or dates? Answer Yes or No: {f_tmp['Details']} "
- answer_tmp = ask_llama(system_message,instruction_message, pipel)
- answer_short = answer_tmp[0]['generated_text'].split("<|eot_id|>\n\nAssistant: ")[-1]
- f_tmp['Details_did'] = answer_short
- except:
- f_tmp['Details_did'] = 'UNKNOWN'
- new_findings.append(f_tmp)
- # save participants finding with scan and brain info
- with open(out_file2, "w") as out:
- json.dump(new_findings, out)
- print('Done :) ')
ask_llama_open.py at commit aa4dce4, under CC-BY-NC-4.0 · at the source
Overview
- Department of Neurology, Beth Israel Deaconess Medical Center (BIDMC), Boston, MA USA
- Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA USA
- Athinoula A. Martinos Center for Biomedical Imaging, Department of Radiology, Massachusetts General Hospital, Charlestown, MA USA
- Stanford University, Palo Alto, USA
- Computer Science & Artificial Intelligence Lab, EECS, MIT, Cambridge, MA USA
- Amazon Web Services, Seattle, USA
- Yale University, New Haven, CT USA
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repository
Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.
bdsp-core/BigBrainImagingDatabase
aa4dce47f3e35c793709874a347f19605fc4f5a6, 24 May 2026Availability: 1 check, the latest on 28 September 2026: the link answers
- 28 September 2026: the link answers
5 files
- ask_llama_category.py, Python, 166 lines, 1 match
- ask_llama_open.py, Python, 233 lines, 2 matches
- mri_sequence_identificat
ion.ipynb , Jupyter, 434 lines, 1 match - LICENSE.txt, License, 413 lines
- README.md, Text, 14 lines
Code availability statement
The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: bdsp-core/
BigBrainImagingDatabase
Read it in the paper: doi.org/10.1038/s41597-026-07421-x.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 3 scripts, each with its path and the digest of its content;
- 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability statement
The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- no repository, dataset or request procedure was recognized in it
Read it in the paper: doi.org/10.1038/s41597-026-07421-x.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 28 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 20 authors, 2 keywords, 10 MeSH terms, 9 funders, 20 references.
Cite
This paper
Maschke, C., Hadar, P. N., Zhang, Y., Li, J., Ganjoo, G., Hoopes, A., Guazzo, A., Gupta, A., Ghanta, M., Nearing, B., Silvers, C. T., Gunapati, B., Thomas, R., Kim, J. A., Mukerji, S. S., Dalca, A., Zafar, S., Lam, A. D., Mignot, E., & Westover, M. B. (2026). The Brain Imaging and Neurophysiology Dataset of large-scale multimodal neural data. Scientific data, 13(1), 1176. https://
BibTeX
@article{maschke2026brai
author = {Maschke, Charlotte and Hadar, Peter N and Zhang, Yicheng and Li, Jian and Ganjoo, Gauri and Hoopes, Andrew and Guazzo, Alessandro and Gupta, Aditya and Ghanta, Manohar and Nearing, Bruce and Silvers, Christine Tsien and Gunapati, Bharath and Thomas, Robert and Kim, Jennifer A and Mukerji, Shibani S and Dalca, Adrian and Zafar, Sahar and Lam, Alice D and Mignot, Emmanuel and Westover, M Brandon},
title = {{The Brain Imaging and Neurophysiology Dataset of large-scale multimodal neural data}},
journal = {Scientific data},
year = {2026},
month = may,
volume = {13},
number = {1},
pages = {1176},
publisher = {Nature Publishing Group},
issn = {2052-4463},
doi = {10.1038/
url = {https://
pmid = {42168237},
pmcid = {PMC13462082}
}
RIS
TY - JOUR
AU - Maschke, Charlotte
AU - Hadar, Peter N
AU - Zhang, Yicheng
AU - Li, Jian
AU - Ganjoo, Gauri
AU - Hoopes, Andrew
AU - Guazzo, Alessandro
AU - Gupta, Aditya
AU - Ghanta, Manohar
AU - Nearing, Bruce
AU - Silvers, Christine Tsien
AU - Gunapati, Bharath
AU - Thomas, Robert
AU - Kim, Jennifer A
AU - Mukerji, Shibani S
AU - Dalca, Adrian
AU - Zafar, Sahar
AU - Lam, Alice D
AU - Mignot, Emmanuel
AU - Westover, M Brandon
TI - The Brain Imaging and Neurophysiology Dataset of large-scale multimodal neural data
T2 - Scientific data
J2 - Sci Data
PY - 2026
DA - 2026/
VL - 13
IS - 1
SP - 1176
SN - 2052-4463
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "The Brain Imaging and Neurophysiology Dataset of large-scale multimodal neural data",
"container-title": "Scientific data",
"author": [
{
"family": "Maschke",
"given": "Charlotte"
},
{
"family": "Hadar",
"given": "Peter N"
},
{
"family": "Zhang",
"given": "Yicheng"
},
{
"family": "Li",
"given": "Jian"
},
{
"family": "Ganjoo",
"given": "Gauri"
},
{
"family": "Hoopes",
"given": "Andrew"
},
{
"family": "Guazzo",
"given": "Alessandro"
},
{
"family": "Gupta",
"given": "Aditya"
},
{
"family": "Ghanta",
"given": "Manohar"
},
{
"family": "Nearing",
"given": "Bruce"
},
{
"family": "Silvers",
"given": "Christine Tsien"
},
{
"family": "Gunapati",
"given": "Bharath"
},
{
"family": "Thomas",
"given": "Robert"
},
{
"family": "Kim",
"given": "Jennifer A"
},
{
"family": "Mukerji",
"given": "Shibani S"
},
{
"family": "Dalca",
"given": "Adrian"
},
{
"family": "Zafar",
"given": "Sahar"
},
{
"family": "Lam",
"given": "Alice D"
},
{
"family": "Mignot",
"given": "Emmanuel"
},
{
"family": "Westover",
"given": "M Brandon"
}
],
"container-title-short":
"volume": "13",
"issue": "1",
"page": "1176",
"DOI": "10.1038/
"PMID": "42168237",
"PMCID": "PMC13462082",
"ISSN": "2052-4463",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
21
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1186/s41747-026-00735-w [code]
- Dual-conditioned diffusion model with anatomical guidance for geometric distortion correction in prostate MRI.Journal: European radiology experimentalIn common: pydicom, Hugging Face Transformers, PyTorch, 2 other tools, structural MRI / diffusion
- [2] doi:10.1111/joa.70203 [code]
- Two-step workflow integrating automatic registration and manual refinement for the accurate alignment of serial histological sections in 3D reconstruction.Journal: Journal of anatomyIn common: pydicom, Hugging Face Transformers, PyTorch, 2 other tools, methods / tools
- [3] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: pydicom, Hugging Face Transformers, PyTorch, 2 other tools
- [4] doi:10.1038/s41597-026-07138-x [code]
- Medical Spine Sagittal MRI Dataset for Segmentation and Foraminal Stenosis detection.Journal: Scientific dataIn common: pydicom, PyTorch, pandas, 1 other tool, methods / tools, structural MRI / diffusion
- [5] doi:10.3389/frai.2026.1771088 [code]
- Few-shot deployment of pretrained MRI transformers in brain imaging tasks.Journal: Frontiers in artificial intelligenceIn common: pydicom, PyTorch, pandas, 1 other tool, methods / tools, structural MRI / diffusion
- [6] doi:10.1093/braincomms/fcag117
- Connectome disruptions after hypoxic-ischaemic injury associate with consciousness disorder severity.Journal: Brain communicationsIn common: structural MRI / diffusion, author M Brandon Westover
- [7] doi:10.1371/journal.pone.0346575 [code]
- Statistically valid explainable black-box machine learning: applications in sex classification across species using brain imaging.Journal: PloS oneIn common: Hugging Face Transformers, PyTorch, pandas, 1 other tool, methods / tools, structural MRI / diffusion
- [8] doi:10.1371/journal.pcbi.1014555 [code]
- Body surface potential driven personalisation of electrophysiological digital twins in hypertrophic cardiomyopathy.Journal: PLoS computational biologyIn common: pydicom, PyTorch, pandas, 1 other tool, structural MRI / diffusion
- [9] doi:10.1080/07853890.2026.2685416 [code]
- Pulmonary and cerebral damage in COVID-19 survivors: is there any association?Journal: Annals of medicineIn common: pydicom, PyTorch, pandas, 1 other tool, structural MRI / diffusion
- [10] doi:10.1371/journal.pone.0354511 [code]
- TokenUNet: A new case for transformers integration in efficient and interpretable 3D UNets for brain imaging segmentation.Journal: PloS oneIn common: Hugging Face Transformers, PyTorch, pandas, 1 other tool, methods / tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 3 scripts, and 4 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:89f2bbd72bb9e253…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
