Recent advances in Multimodal Large Language Models (MLLMs) have created new opportunities for intelligent video analysis by enabling semantic reasoning across visual and textual modalities. This study presents a novel Cross-Modal Knowledge Mining (CMKM) framework for automated video scene understanding and event detection. The proposed framework integrates visual feature extraction, semantic knowledge generation, temporal event saliency estimation, and multimodal fusion to establish bidirectional interactions between video content and language-based representations. By leveraging the complementary strengths of visual and semantic information, the framework enhances contextual understanding and improves event recognition performance. Extensive experiments conducted on multiple benchmark video datasets demonstrate the effectiveness and robustness of the proposed approach under supervised, few-shot, and zero-shot learning settings. The results indicate that cross-modal knowledge mining significantly improves scene interpretation, event detection accuracy, and model generalization, highlighting the potential of MLLMs for next-generation video intelligence systems.
Jalal et al. (Sat,) studied this question.