PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 20260 citationsOpen Access

Interrater Reliability Comparisons with Generalizability Theory and Structural Equation Modeling

View Full Paper
HFHolmes FinchBFBrian FrenchJIJason C. Immekus

Key Points

  • This research aims to refine the assessment of interrater reliability using advanced statistical methodologies.
  • Conducted a simulation study comparing generalizability theory (GT) and structural equation modeling (SEM) methods against the W statistic.
  • Tested differences in reliability coefficients across various groups and conditions.
  • Evaluated Type I error control and statistical power in moderate to large sample sizes.
  • The proposed GT and SEM method consistently outperformed the W statistic, leading to better Type I error control.
  • The method showed enhanced statistical power, especially when rater agreement variance existed across groups.

Abstract

Interrater reliability is a critical aspect of measurement quality, particularly in assessments that rely on subjective judgment. However, interrater reliability estimates vary, and such variability can introduce bias or reduce the accuracy of observed scores, especially when comparing across groups or conditions. Understanding and accounting for these differences is essential when interpreting reliability in applied settings such as education, psychology, and performance evaluation. This study addresses the need for more nuanced approaches to evaluating interrater reliability across groups. Specifically, in this study, we examine generalizability theory (GT) and structural equation modeling (SEM) that enable direct testing of differences in reliability coefficients across groups. A simulation study compared a proposed method grounded in GT and SEM to the W statistic for reliability coefficient comparisons. Results demonstrate that the proposed method consistently outperforms the W statistic in terms of both Type I error control and statistical power, particularly when sample sizes are moderate to large or when variance in rater agreement exists across groups. These findings underscore the importance of explicitly modeling differences in interrater reliability and provide researchers with a more robust tool for evaluating the consistency of ratings across diverse contexts and populations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Finch et al. (2026) studied this question.

synapsesocial.com/papers/69b25adb96eeacc4fcec8fcfhttps://doi.org/10.3390/psycholint8010019
Ask AI
Helpful
Bookmark
Share
View Full Paper