This paper presents the Social Analytics API, a scalable multi-modal machine learning system designed for pre-publication analysis of short-form video content. The system integrates visual embeddings from CLIP with textual embeddings derived from video metadata (titles, descriptions, and tags) to form a fused representation for classification and clustering. A LightGBM classifier trained on the fused feature space achieves 88.0% macro-average accuracy, outperforming the visual-only baseline by 11.4 percentage points. The system further introduces sub-niche discovery using K-Means clustering, generating 63 semantically coherent clusters across 12 macro categories. Additionally, the architecture incorporates structural video analysis modules and a live market validation layer using YouTube Data API v3 to compute metrics such as Supply Factor and Trend Alignment Score. The proposed approach demonstrates the effectiveness of multi-modal feature fusion for improving video intelligence systems and enabling data-driven content optimization before publication.
Jaskirat Singh (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: