Advantage of grading classification using volumetric artificial intelligence for periventricular hyperintensity and deep subcortical white matter hyperintensity
Abstract
Abstract We developed and validated an artificial intelligence (AI) algorithm for the automated grading of periventricular hyperintensity (PVH) and deep subcortical white matter hyperintensity (DWMH) using magnetic resonance imaging. Overall, 246 patients were evaluated, with 137 and 109 allocated to the training and testing groups, respectively. AI-predicted grading according to the Fazekas scale was compared with expert assessments using accuracy, F1-score, and mean absolute error. Inter-rater agreement was evaluated using Fleiss’ kappa to assess consistency among human raters and Cohen’s kappa to measure agreement between the AI and individual human raters. The AI demonstrated superior multi-class accuracy in PVH classification compared with the human expert, achieving an accuracy of 0.798 versus 0.743. In DWMH classification, the AI outperformed the expert specifically in distinguishing Fazekas 0/1/2 from the 3 classification, achieving an accuracy of 0.954 compared with the expert’s 0.927. Inter-rater agreement analysis showed that for PVH and DWMH, the AI achieved “good agreement” with human raters. For PVH, the AI’s agreement exceeded the human inter-rater agreement. The developed AI also exhibited lower variability in volume ratio distribution within the same grade compared with human raters. The developed AI algorithm effectively distinguished between PVH and DWMH, achieving accuracy comparable to human performance.
Article Details
Authors (14)
Masashi Kuwabara
Fusao Ikawa
Shinji Nakazawa
Saori Koshino
Daizo Ishii
Hiroshi Kondo
Takeshi Hara
Shingo Matsuda
Yuyo Maeda
Shiyuki Maeyama
Yoshinobu Seo
Jinichi Sasanuma
Kimito Kondo
Nobutaka Horie