In response to the inefficiency, difficulty in unifying standards, and difficulty in integrating multimodal data such as text, images, and audio caused by the long-term reliance on manual operations in the classification and storage process of traditional archival resources, this study designs and implements an automated processing solution. This solution integrates multimodal feature extraction and machine learning classification methods, and constructs a complete technical process including data preprocessing, feature construction, model optimization, and storage adaptation. To test the effectiveness of machine learning in automated archive processing, this study compared the performance of three types of models: support vector machine, random forest, and Transformer. The experimental results show that the Transformer based multimodal fusion model performs the best in archive text, image, and audio classification tasks, with a comprehensive accuracy of 92.3%, which is 15.1% higher than support vector machines and 8.7% higher than random forests. In addition, by compressing the classification model and improving the storage strategy, the required storage space can be reduced by 40%, thereby better meeting the practical needs of archive management systems for efficiency, accuracy, and security. This study provides a feasible and effective technical path for the intelligent transformation of archive resource management.
Zhan Gao (Thu,) studied this question.