This article explores the convergent validity—or, more precisely, the inter-rater consistency—of popular generative artificial intelligence (AI) models in evaluating the quality of project management (PM) software. This study employed three prominent generative AI models—Gemini, CoPilot on Edge, and DeepSeek—to independently assign rating scores (1-10) to a curated list of 50 top-tier PM software systems/tools. The evaluation was structured around eight critical dimensions. Statistical analyses were conducted to assess the consistency of the AI-generated ratings. These included descriptive statistics to measure rating discrimination, analysis of mean absolute differences, paired-samples t -tests to identify erratic and systematic rating biases between AI model pairs, and Cronbach’s alpha to determine overall inter-rater consistency for each dimension. The results reveal a high degree of inter-rater consistency across all eight dimensions, a finding juxtaposed with the presence of significant systematic biases between the models. Cronbach’s alpha coefficients were found to be acceptable for all dimensions ( α = .782 to .897), indicating that the models consistently rate the PM software in a similar order. However, the t -tests confirmed that each model possessed a distinct "rating personality," consistently scoring higher or lower than its counterparts. These findings suggest that while generative AI is a surprisingly trustworthy measure for rating PM software roughly in a particular order, decision-makers must account for the distinct rating biases of each model. Relying on a single model for absolute rating scores is ill-advised, but leveraging multiple models to establish a consensus on relative quality is a viable and powerful new approach for PM software evaluation.
Victor K.Y. Chan (2026) studied this question.