The Analytic Hierarchy Process (AHP) is widely used in Multi-Criteria Decision Analysis (MCDA), yet its strong reliance on expert judgment constrains its scalability and may introduce variability in weighting outcomes, particularly in high-stakes applications such as wildfire risk assessment. In this study, we investigate how Large Language Models (LLMs) can function as decision-support agents in an AHP-style hierarchical evaluation task derived from validated wildfire literature. Based on this structure, four representative LLM-assisted strategies are examined: Direct LLM Scoring (DLS), Multi-Model Debate Scoring (MDS), Full-Document Prompting (FDP), and Indicator-Guided Prompting (IGP). To evaluate their effectiveness, we benchmark LLM-generated rankings against expert-defined ground truth across 16 sub-criteria. Using the mean correlation coefficient R as the key evaluation metric, with reported values expressed as mean ± standard deviation across models: DLS shows no correlation with expert rankings (R = 0.009 ± 0.070), MDS yields marginal gains (R = 0.181), and FDP remains unstable (R = 0.081 ± 0.189). By contrast, IGP, which incorporates retrieval-informed structured prompting, shows the highest agreement with the expert reference among the four compared strategies (R = 0.598 ± 0.065), suggesting that structured contextual guidance may improve the performance of LLM-assisted weighting within the evaluated benchmark. This study suggests that, within the evaluated wildfire benchmark and the tested set of hosted LLMs, LLMs may serve as useful decision-support tools in MCDA tasks when guided by structured inputs or coordinated through multi-agent mechanisms. The proposed framework provides an interpretable basis for exploring LLM-assisted risk evaluation in the present wildfire benchmark, while further validation is needed before extending it to other environmental or safety-critical contexts.
Cheng et al. (Thu,) studied this question.