Key points are not available for this paper at this time.
Safety-critical perception systems depend on rare, stochastic events that are difficult to collect at scale due to adverse conditions and ethical constraints. Conventional augmentation operates at the pixel level and cannot introduce new object interactions or high-risk configurations, leaving datasets sparse at the semantic-event level. Recent generative models enable photorealistic synthesis from natural language prompts, yet most lack explicit semantic control, systematic evaluation, and iterative improvement mechanisms. This paper proposes an agentic ontology-guided framework for rare-event synthetic image generation, using wildlife–traffic interaction as a safety-critical case study. The framework integrates a formal semantic ontology, autonomous generation agents, LMM-based evaluation agents, and a closed-loop refinement and regeneration mechanism, enabling controlled comparison between reference-based and referenceless generation across OpenAI, Gemini, and Grok models. Evaluation combines objective no-reference image quality metrics, LMM-as-a-Judge, a Panel of LMM Evaluators, human ratings, and downstream zero-shot object detection. Among referenceless models, Gemini 3 Pro Image consistently leads on perceptual criteria while OpenAI-generated images achieve particularly strong recall, mAP, and false-negative reduction, especially under GDINO. These findings demonstrate that comprehensive evaluation of rare-event synthetic data requires both perceptual quality metrics and task-aware detectability criteria, and that the proposed framework supports this multi-dimensional assessment within a unified agentic pipeline.
Alaa Khamis (2026) studied this question.