The complex environment of electrical work sites presents hazards that are diverse in form, easily concealed, and difficult to distinguish from their surroundings. Due to poor model generalization, most traditional visual recognition methods are prone to errors and cannot meet the current safety management needs in electrical work. This paper presents a novel framework for hazard identification that integrates chain-of-thought reasoning and self-verification mechanisms within a visual-language large model (VLLM) to enhance accuracy. First, typical hazard scenario data for crane operation and escalator work areas were collected. The Janus-Pro VLLM model was selected as the base model for hazard identification. Then, designing a chain-of-thought enhanced the model’s capacity to identify critical information, including the status of crane stabilizers and the zones where personnel are located. Simultaneously, a self-verification module was designed. It leveraged the multimodal comprehension capabilities of the VLLM to self-check the identification results, outputting confidence scores and justifications to mitigate model hallucination. The experimental results show that integrating the self-verification method significantly improves hazard identification accuracy, with average increases of 2.55% in crane operations and 4.35% in escalator scenarios. Compared with YOLOv8s and D-FINE, the proposed framework achieves higher accuracy, reaching up to 96.3% in crane personnel intrusion detection, and a recall of 95.6%. It outperforms small models by 8.1–13.8% in key metrics without relying on massive labeled data, providing crucial technical support for power operation hazard identification.
Gao et al. (Wed,) studied this question.