Abstract Natural language processing (NLP) models can classify unstructured social media content into structured data suitable for generating public health intelligence. NLP models, however, can vary in performance, given differences in architecture and sensitivities to temporal drift. Our objective was measuring the performance of different NLP models for identifying topics and sentiment of social media posts related to COVID-19 within a framework that integrated model outputs and human-labelled posts. Posts were obtained from Twitter (Canada; December 2020). Tweets were classified by public health measure (PHM) topic (mask-wearing, vaccination, social distancing, others) using five crowdsourcing annotators, a keyword-based filter (KB) and a large language model (LLM). Sentiment (in support, neutral, against) was categorized using the same crowdsourcing annotators and LLM, and a neural network sentiment analysis (NSA) model. Sensitivity and specificity of the NLP models were estimated using hierarchical Bayesian latent class models. A total of 443 and 801 tweets were assessed for topic and sentiment classification, respectively. KB and LLM models had similar specificity and sensitivity for all topics, except for LLM having higher sensitivity for vaccination (0.88). The NSA model yielded sensitivity 0.50 for classifying “in support” and “against” sentiments, while the LLM had sensitivity 0.60 for all sentiments. Deployment of NLP models for public health applications requires critical understanding of model performance. This study demonstrates an approach to estimate the posterior distribution of the probability, adjusted for misclassification error, of including topic/sentiment for each social media post and derived classifier accuracy metrics for the tests.
Denis-Robichaud et al. (Tue,) studied this question.