SUMMARY This study examines whether large language models can predict critical audit matters (CAMs) from management discussion and analysis (MD&A) disclosures in 10-K filings. We tested four models, OpenAI’s O3-Pro and O3-Mini High, Anthropic’s Claude Sonnet 4, and Google’s Gemini 2.5 Flash, using zero-shot prompts against auditor-identified CAMs across 141 firm-year observations in technology, finance, consumer discretionary, and healthcare. We benchmarked predictive capability matching auditor CAMs and generative capacity, identifying audit-relevant risk matters from MD&A. O3-Mini High demonstrated the strongest predictive performance, whereas Gemini 2.5 Flash underperformed in both areas. Technology achieved the highest prediction accuracy, whereas healthcare ranked last. We also examined how management tone and algorithmic bias jointly affect prediction outcomes. These findings position artificial intelligence (AI) as a valuable screening tool to support auditor planning, while reinforcing that professional judgment remains essential to final CAM determination. Data Availability: Our data will be made available upon request.
Akpan et al. (Fri,) studied this question.