4.7 Article

Reproducibility of radiomics quality score: an intra- and inter-rater reliability study

期刊

EUROPEAN RADIOLOGY
卷 -, 期 -, 页码 -

出版社

SPRINGER
DOI: 10.1007/s00330-023-10217-x

关键词

Reproducibility of results; Artificial intelligence; Radiomics; Inter-observer variability; Intra-observer variability

向作者/读者索取更多资源

This study investigated the reliability of the total radiomics quality score (RQS) and the reproducibility of individual RQS items' score in a large multireader study. The results showed low inter-rater reliability for total RQS and moderate to good intra-rater reliability for individual RQS items' score. There is a need for a robust and reproducible assessment method to improve the quality of radiomics research.
Objectives To investigate the intra- and inter-rater reliability of the total radiomics quality score (RQS) and the reproducibility of individual RQS items' score in a large multireader study. Methods Nine raters with different backgrounds were randomly assigned to three groups based on their proficiency with RQS utilization: Groups 1 and 2 represented the inter-rater reliability groups with or without prior training in RQS, respectively; group 3 represented the intra-rater reliability group. Thirty-three original research papers on radiomics were evaluated by raters of groups 1 and 2. Of the 33 papers, 17 were evaluated twice with an interval of 1 month by raters of group 3. Intraclass coefficient (ICC) for continuous variables, and Fleiss' and Cohen's kappa (k) statistics for categorical variables were used. Results The inter-rater reliability was poor to moderate for total RQS (ICC 0.30-055, p < 0.001) and very low to good for item's reproducibility (k - 0.12 to 0.75) within groups 1 and 2 for both inexperienced and experienced raters. The intra-rater reliability for total RQS was moderate for the less experienced rater ( ICC 0.522, p = 0.009), whereas experienced raters showed excellent intra-rater reliability (ICC 0.91-0.99, p < 0.001) between the first and second read. Intra-rater reliability on RQS items' score reproducibility was higher and most of the items had moderate to good intra-rater reliability (k - 0.40 to 1). Conclusions Reproducibility of the total RQS and the score of individual RQS items is low. There is a need for a robust and reproducible assessment method to assess the quality of radiomics research. Clinical relevance statement There is a need for reproducible scoring systems to improve quality of radiomics research and consecutively close the translational gap between research and clinical implementation.

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

4.7
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据