Balancing privacy and performance: the impact of facial defacing on AI in medical imaging.
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Markdown · 76 lines · 4.1 KB · no license
- # Balancing Privacy and Performance: The Impact of Facial Defacing on AI in Medical Imaging
- <img src="Pics/Overall_pipeline.png" align="middle" width="75%">
- The 2025 NIH Data Management and Sharing policy mandates facial anonymization for publicly shared medical imaging data, prompting concerns about potential degradation of AI model performance. We systematically evaluated three representative defacing algorithms, two invasive (QuickShear, Deface) and one geometry-preserving, non-invasive method (Reface), across MRI and CT datasets from 600 subjects. Model performance was assessed across complementary clinical tasks: (1) brain segmentation and Evans ratio measurement in MRI of normal pressure hydrocephalus (NPH), (2) representative slice selection and diagnostic reasoning for brain tumor MRI using vision–language models (VLMs), and (3) automated report generation from emergency department CT head scans via VLMs. Invasive defacing substantially impaired segmentation accuracy, diagnostic precision, and report fidelity, whereas the non-invasive Reface approach preserved model performance with minimal deviation from original data. These findings quantify the privacy–utility trade-off introduced by facial defacing and underscore the importance of non-invasive, geometry-preserving anonymization techniques to ensure both patient privacy and AI model reproducibility under the emerging NIH data-sharing mandate.
- ## Datasets Description
- This retrospective study was approved by the Institutional Review Boards of Johns Hopkins University (IRB00318113 for the NPH MRI dataset, IRB00452757 for the Brain Tumor MRI dataset, and IRB00424745 for the ED head CT dataset) with a waiver of informed consent. A data transfer agreement with the University of Pennsylvania (identification number: 73481-00) was also established for the Brain Tumor MRI dataset.
- ### Data Structure
- | Type | No. (ALL) | No. (Include) | Format |
- | --------------------------| ------------| ------------- | -------|
- | NPH MRI Dataset | 534 | 200 | DICOM |
- | Brain Tumor MRI Dataset | 2453 | 200 | NiFTI |
- | ED CT Head Dataset | 33,128 | 200 | DICOM |
- ## Deface methods and implementation details
- ### Quickshear
- **Version:** `v1.1.0` (May 23, 2017)
- **Repository:** [nipy/quickshear](https://github.com/nipy/quickshear)
- Quickshear removes facial features by defining a **shearing plane** separating the brain from the face and setting all voxels anterior to the plane to zero.
- The plane is derived from a brain mask (e.g., via BET) and propagated across all sagittal slices.
- A tunable buffer controls how aggressively the face is removed.
- #### Installation
- ```
- pip install quickshear
- # or clone
- git clone https://github.com/nipy/quickshear
- ```
- ### Deface (Deep Learning)
- **Version:** ver.0.2 (latest commit, November 2020)
- **Repository:** yeonuk-Jeong/Defacer
- Defacer is the first open-source deep learning–based anonymization tool.
- It uses a 3D attention-gated U-Net trained on 240 whole-head MRIs from ADNI to detect and modify facial features (eyes, ears, nose, mouth) while preserving the brain.
- Validated on 100 OASIS MRIs.
- #### Installation
- ```
- git clone https://github.com/yeonuk-Jeong/Defacer
- cd Defacer
- pip install tensorflow==1.14 keras==2.2.4 nibabel
- ```
- ### Reface (mri_reface)
- **Version:** v0.2, v0.3, and latest v0.3.5 (Dec 2024)
- **Repository:** mri_reface on NITRC
- Reface replaces (rather than removes) facial voxels.
- An average face is registered to the subject using ANTs or NiftyReg, and the warped face is blended into the image with bias correction and intensity harmonization.
- This preserves brain structures while anonymizing facial identity.
- #### Installation
- ```
- wget https://www.nitrc.org/frs/download.php/13678/mri_reface_v0.3.5_Linux.zip
- unzip mri_reface_v0.3.5_Linux.zip
- export PATH=$PATH:$PWD/mri_reface_v0.3.5
- ```
- ## Downstream tasks
- ### Brain Sementations
- **Freesurfer:** [Freesurfer](https://github.com/freesurfer/freesurfer)
- **SLANT:** [SLANT](https://github.com/MASILab/SLANTbrainSeg)
- ### Representative slides selection
- **Vote-MI:** [Freesurfer](https://github.com/YuliWanghust/BrainMD_NIPS)
README.md at commit 43c816f, no license · at the source
Overview
- Department of Radiology, University of Colorado Anschutz Medical Campus, Aurora, CO, USA
- Department of Radiology and Radiological Science, Johns Hopkins University School of Medicine, Baltimore, MD, USA
- Department of Neurology, Second Xiangya Hospital of Central South University, Changsha, Hunan, China
- Department of Computer Science, Johns Hopkins University, Baltimore, MD, USA
- Department of Radiology, Second Xiangya Hospital of Central South University, Changsha, Hunan, China
- Department of Pathology and Laboratory Medicine, University of Pennsylvania, Philadelphia, PA, USA
- School of Humanities, Central South University, Changsha, China
- Department of Neurosurgery, Massachusetts General Hospital, Boston, MA, USA
- Department of Diagnostic Imaging, Brown University Health, Providence, RI, USA
- Department of Radiology, Beijing Tiantan Hospital, Capital Medical University, Beijing, China
Abstract
Background: Recent NIH Data Management and Sharing (DMS) policy updates and NIH controlled-access data security requirements have increased attention to facial anonymization and controlled-access handling of shared head imaging data. This is particularly relevant for datasets submitted to or hosted by the Cancer Imaging Archive (TCIA), where NCI Cancer Imaging Program/
Methods: We systematically evaluated three representative defacing algorithms, two invasive (QuickShear and Py-Deface) and one less destructive, facial replacement (mri_reface), across MRI and CT datasets from 600 subjects spanning three institutions. Model performance was assessed on three clinically relevant applications: (1) brain segmentation and Evans ratio biomarker quantification in normal pressure hydrocephalus (NPH) MRI using SLANT and FreeSurfer; (2) representative-slice selection and diagnostic reasoning for brain tumour MRI using vision-language models (VLMs); and (3) automated emergency head CT report generation using a fine-tuned Otter-based vision-language model. Each method’s impact was quantified using Dice similarity, correlation metrics, reasoning accuracy, and natural-language generation scores (BLEU, METEOR, ROUGE, CIDEr).
Findings: Invasive algorithms caused significant degradation across all tasks. QuickShear reduced mean Dice scores by up to 9% and introduced 14–19% failure rates during quality control, while PyDeface induced smaller but measurable performance losses. mri_reface maintained 100% success without any failures and achieved segmentation, diagnostic, and report-generation accuracy within 3–5% of the original data. Evans ratio distributions remained statistically consistent between mri_reface and original images (p > 0.05), whereas invasive methods introduced broader variance. Across all VLM tasks, mri_reface preserved high correlation with radiologist-selected slices (r = 0.979) and stable report-generation quality (BLEU-4 = 0.11 ± 0.06 vs. 0.12 ± 0.07 for original).
Interpretation: Facial anonymization introduces a measurable privacy-utility trade-off that must be explicitly considered in the design of AI-ready medical imaging datasets. Invasive defacing compromises geometric and statistical integrity, reducing downstream model accuracy even outside facial regions. Facial replacement anonymization methods, such as mri_reface, effectively reconcile patient privacy with reproducibility, offering a practical path to NIH-compliant open data. Future regulatory and institutional policies should integrate quantitative privacy-utility assessment and mandate transparent reporting of anonymization pipelines to ensure that shared imaging data remain both ethically safe and scientifically valid under emerging digital health frameworks.
Funding: This work was partially supported by the 10.13039/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above.
YuliWanghust/deface_nih
43c816f3cb0022c041b743643c7db4c2c1f3b9fb, 24 October 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
1 file
- README.md, Text, 76 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 0 scripts, each with its path and the digest of its content;
- no match between paragraphs and code yet;
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data sharing statement
The datasets used in this study are not publicly available due to institutional, regulatory, and patient privacy restrictions. Reasonable requests for data access may be directed to the corresponding author and will be considered subject to IRB approval and institutional data-use agreements. The custom code and scripts are available at the GitHub link available at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Authors: added Cheng Ting Lin (0000-0003-0275-8037); removed Cheng Ting Lin
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 19 authors, 5 keywords, 9 MeSH terms, 4 funders, 25 references.
Cite
This paper
Wang, Y., Dai, Y., Guan, H., Li, C.-Y., Vora, M., Ujano, D., Nandi, A., Wu, J., Zhang, P., Honce, J. M., Zhu, C., Sair, H. I., Balaj, L., Jiao, Z., Liu, Y., Lin, C. T., Kamel, I., Yang, L., & Bai, H. (2026). Balancing privacy and performance: the impact of facial defacing on AI in medical imaging. EBioMedicine, 131, 106457. https://
BibTeX
@article{wang2026balanci
author = {Wang, Yuli and Dai, Yuwei and Guan, Haoyue and Li, Cheng-Yi and Vora, Maulik and Ujano, Danica and Nandi, Ayon and Wu, Jing and Zhang, Paul and Honce, Justin M. and Zhu, Chengzhang and Sair, Haris I. and Balaj, Leonora and Jiao, Zhicheng and Liu, Yaou and Lin, Cheng Ting and Kamel, Ihab and Yang, Li and Bai, Harrison},
title = {{Balancing privacy and performance: the impact of facial defacing on AI in medical imaging}},
journal = {EBioMedicine},
year = {2026},
month = aug,
volume = {131},
pages = {106457},
publisher = {Elsevier},
issn = {2352-3964},
doi = {10.1016/
url = {https://
pmid = {42667924},
pmcid = {PMC13545401}
}
RIS
TY - JOUR
AU - Wang, Yuli
AU - Dai, Yuwei
AU - Guan, Haoyue
AU - Li, Cheng-Yi
AU - Vora, Maulik
AU - Ujano, Danica
AU - Nandi, Ayon
AU - Wu, Jing
AU - Zhang, Paul
AU - Honce, Justin M.
AU - Zhu, Chengzhang
AU - Sair, Haris I.
AU - Balaj, Leonora
AU - Jiao, Zhicheng
AU - Liu, Yaou
AU - Lin, Cheng Ting
AU - Kamel, Ihab
AU - Yang, Li
AU - Bai, Harrison
TI - Balancing privacy and performance: the impact of facial defacing on AI in medical imaging
T2 - EBioMedicine
J2 - EBioMedicine
PY - 2026
DA - 2026/
VL - 131
SP - 106457
SN - 2352-3964
PB - Elsevier
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "Balancing privacy and performance: the impact of facial defacing on AI in medical imaging",
"container-title": "EBioMedicine",
"author": [
{
"family": "Wang",
"given": "Yuli"
},
{
"family": "Dai",
"given": "Yuwei"
},
{
"family": "Guan",
"given": "Haoyue"
},
{
"family": "Li",
"given": "Cheng-Yi"
},
{
"family": "Vora",
"given": "Maulik"
},
{
"family": "Ujano",
"given": "Danica"
},
{
"family": "Nandi",
"given": "Ayon"
},
{
"family": "Wu",
"given": "Jing"
},
{
"family": "Zhang",
"given": "Paul"
},
{
"family": "Honce",
"given": "Justin M."
},
{
"family": "Zhu",
"given": "Chengzhang"
},
{
"family": "Sair",
"given": "Haris I."
},
{
"family": "Balaj",
"given": "Leonora"
},
{
"family": "Jiao",
"given": "Zhicheng"
},
{
"family": "Liu",
"given": "Yaou"
},
{
"family": "Lin",
"given": "Cheng Ting"
},
{
"family": "Kamel",
"given": "Ihab"
},
{
"family": "Yang",
"given": "Li"
},
{
"family": "Bai",
"given": "Harrison"
}
],
"container-title-short":
"volume": "131",
"page": "106457",
"DOI": "10.1016/
"PMID": "42667924",
"PMCID": "PMC13545401",
"ISSN": "2352-3964",
"publisher": "Elsevier",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
29
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1002/mrm.70380 [code]
- The Impact and Reliability of Tissue Segmentation on In Vivo Magnetic Resonance Spectroscopy Metabolite Quantification.Journal: Magnetic resonance in medicineIn common: structural MRI / diffusion, 2 references
- [2] doi:10.1162/imag.a.1183 [code]
- Learning-based segmentation of diffusion-weighted MR images with arbitrary &
lt;i& gt;q& lt;/ i& gt;-space samplings. Journal: Imaging neuroscience (Cambridge, Mass.)In common: structural MRI / diffusion, 2 references - [3] doi:10.7554/elife.108109 [code]
- Multimodal MRI marker of cognition explains the association between cognition and mental health in the UK Biobank.Journal: eLifeIn common: structural MRI / diffusion, 2 references
- [4] doi:10.1002/dneu.70019 [code]
- A Longitudinal Study of Children's Hippocampal Development: Investigating Maternal Physical Activity, Depression, and Education.Journal: Developmental neurobiologyIn common: structural MRI / diffusion, 1 reference
- [5] doi:10.1162/imag.a.1337 [code]
- Data quality biases normative models derived from fetal brain MRI.Journal: Imaging neuroscience (Cambridge, Mass.)In common: structural MRI / diffusion, 2 references
- [6] doi:10.1038/s41598-026-55397-w [code]
- Fast surface reconstruction of human brain MRI: benchmarking deep-learning based morphometry tools.Journal: Scientific reportsIn common: structural MRI / diffusion, 1 reference
- [7] doi:10.1038/s41538-026-00914-4
- Green tea catechin EGCG attenuates hippocampal atrophy and cognitive impairment in obesity via autophagy signaling.Journal: NPJ science of foodIn common: structural MRI / diffusion, other condition, 1 reference
- [8] doi:10.1038/s41586-026-10345-6 [code]
- Population-scale repeat expansions elucidate disease risk and brain atrophy.Journal: NatureIn common: structural MRI / diffusion, other condition, 1 reference
- [9] doi:10.64898/2026.03.10.710798 [code]
- Exploring Links between Brain Image-Derived Phenotypes and Accelerometer-Measured Physical Activity in the UK BiobankJournal: bioRxiv (preprint)In common: structural MRI / diffusion, other condition, 1 reference
- [10] doi:10.1093/cercor/bhag125 [code]
- Heritable and experience-dependent cortical traits of reading ability.Journal: Cerebral cortex (New York, N.Y. : 1991)In common: other condition, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 0 scripts, and 0 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:5ee2105e30b550a2…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
