PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 26, 2026Procedia Computer Science0 citationsOpen Access

The Convergent Validity of Project Management Software Evaluation by Generative Artificial Intelligence Models

View Full Paper
VCVictor K.Y. Chan

Key Points

  • This research aims to assess the inter-rater consistency of generative AI models in evaluating project management software.
  • Evaluated 50 PM software systems using three generative AI models: Gemini, CoPilot on Edge, DeepSeek.
  • Ratings were assigned on an ordinal scale from 1 to 10 across eight critical dimensions.
  • Statistical measures included descriptive statistics, paired-samples t-tests, and Cronbach’s alpha to assess consistency and biases.
  • High inter-rater consistency was found across all eight dimensions with Cronbach’s alpha coefficients ranging from .782 to .897.
  • Significant systematic biases were identified, revealing each AI model had a unique rating tendency.
  • Using multiple AI models for evaluation is recommended to mitigate biases and establish a consensus on software quality.

Abstract

This article explores the convergent validity—or, more precisely, the inter-rater consistency—of popular generative artificial intelligence (AI) models in evaluating the quality of project management (PM) software. This study employed three prominent generative AI models—Gemini, CoPilot on Edge, and DeepSeek—to independently assign rating scores (1-10) to a curated list of 50 top-tier PM software systems/tools. The evaluation was structured around eight critical dimensions. Statistical analyses were conducted to assess the consistency of the AI-generated ratings. These included descriptive statistics to measure rating discrimination, analysis of mean absolute differences, paired-samples t -tests to identify erratic and systematic rating biases between AI model pairs, and Cronbach’s alpha to determine overall inter-rater consistency for each dimension. The results reveal a high degree of inter-rater consistency across all eight dimensions, a finding juxtaposed with the presence of significant systematic biases between the models. Cronbach’s alpha coefficients were found to be acceptable for all dimensions ( α = .782 to .897), indicating that the models consistently rate the PM software in a similar order. However, the t -tests confirmed that each model possessed a distinct "rating personality," consistently scoring higher or lower than its counterparts. These findings suggest that while generative AI is a surprisingly trustworthy measure for rating PM software roughly in a particular order, decision-makers must account for the distinct rating biases of each model. Relying on a single model for absolute rating scores is ill-advised, but leveraging multiple models to establish a consensus on relative quality is a viable and powerful new approach for PM software evaluation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Victor K.Y. Chan (2026) studied this question.

synapsesocial.com/papers/69c4cd98fdc3bde44891a1c0https://doi.org/10.1016/j.procs.2026.03.193
Ask AI
Helpful
Bookmark
Share
View Full Paper