Visual fidelity–driven quality assessment of medical image translation
Abstract
Abstract Automated and reliable image quality assessment (IQA) is essential for safe use of medical image synthesis in critical applications like adaptive radiotherapy, treatment planning, or missing-modality reconstruction, where unnoticed generative artifacts may adversely affect outcomes. We evaluated image-to-image translation quality by coupling large-scale visual quality assessment with explainable automated IQA modeling. Adversarial diffusion-based framework, SynDiff, was applied to four cross-modality synthesis tasks, including three inter-MR and a CBCT-to-CT translation. Using four-fold cross-validation, ten reference-based and eight no-reference IQA metrics were computed for all synthesized images. Visual IQA ratings were independently collected from thirteen raters using predetermined protocol and specialized image viewer enabling blinded, randomized six-point Likert scoring. Auto-Sklearn was employed to learn ensemble regression models mapping IQA metrics to visual consensus ratings, with separate models trained on reference-based and no-reference metrics. The models closely reproduced distribution and ordering of ratings, typically within ±0.5 Likert points. Reference-based models achieved higher agreement with visual ratings than no-reference models ( $$R^2$$ 0.75 vs. 0.59, resp.), although the latter remained unbiased and informative. Explainability analyses indicated that metrics quantifying structural similarity or fidelity (e.g., anatomical boundary preservation) and intensity-based contrast relationships between tissues were the strongest predictors. Overall, the results demonstrate that ensemble regression models can provide transparent, scalable, and clinically meaningful quality control for generative medical imaging.
Article Details
Authors (3)
Žiga Bizjak
Jan Žagar
Žiga Špiclin