Abstract Peer review of research products suffers from poor inter‐rater reliability. Few studies examine whether this limitation generalizes to case reports. We conducted a cross‐sectional analysis of peer reviews of clinical vignette abstracts submitted to a national hospitalist meeting in 2024 and 2025. Three randomly assigned reviewers scored each vignette on a 1–10 scale. We analyzed variation in scores across abstracts and reviewers and estimated inter‐rater reliability via intraclass correlation coefficient (ICC). Two hundred twenty‐one reviewers evaluated 1630 abstracts in 2024–2025. Abstract scores varied substantially: 384/1630 (23.6%) abstracts had a difference of 4 or more points (>2 standard deviations) between highest and lowest reviewer scores. Scores varied by reviewer: 2024 reviewer‐level mean scores ranged 4.27–8.47 (standard deviation (SD): 0.70–2.80); 2025 scores ranged 4.06–8.59 (SD: 0.62‐2.69). Inter‐rater reliability was poor (ICC: 0.37). Adjusting final scores based on reviewer scoring tendencies changed the accept/reject category for 183 (11.2%) abstracts, suggesting opportunities for quality improvement.
Stephens et al. (Thu,) studied this question.