Ship collisions pose substantial risks to maritime safety, causing vessel damage, casualties, and environmental impacts. Efficient extraction and analysis of key navigational and causal information from accident reports are important for risk assessment and decision support. This study proposes a framework for synthetic data generation, DistilBERT-based named entity recognition, and structured dataset construction for ship collision accidents. Using a template-based method, 56,000 annotated sentences were generated, covering navigational elements and causal factor trigger phrases. The fine-tuned DistilBERT model showed good performance on both synthetic and real accident reports. Statistical and co-occurrence analyses further indicated that failure to maintain proper lookout, failure to take effective evasive action, and failure to maintain safe speed were the main contributing factors across different environments and accident severity levels. Based on the extraction results, a standardized structured dataset was constructed to support subsequent causal analysis, dynamic risk modeling, and collision risk prediction. The study shows that combining template-based data synthesis with Transformer-based named entity recognition is a feasible approach for extracting information from maritime accident reports and transforming unstructured text into structured datasets.
Zhang et al. (Thu,) studied this question.