In response to the issue of insufficient prediction accuracy due to low proportions and imbalanced categories of electricity-sensitive users in power user data, this paper proposes an imbalanced data classification model based on the fusion of Imbalance-XGBoost and LightGBM algorithms. First, a multi-source feature set is constructed via one-hot encoding, time windows, TF-IDF vectorization., and other methods. Second, high-dimensional sparse text features are screened according to feature importance using LightGBM. Then, an Imbalance-XGBoost-based weighted cross-entropy loss function with category weights and regularization optimizes minority class learning, while LightGBM’s histogram binning and leaf optimization boost computational efficiency. Finally, a Stacking ensemble strategy is adopted, integrating the prediction results of ridge regression models with two base learners to construct a fusion model. Experiments based on electricity usage data from a power company show that the F1 score of the fusion model reaches 89.7%, respectively, effectively reducing the misclassifi-cation rate. The study demonstrates that the proposed method enhances the recognition capability of minority class samples while balancing model efficiency and generalization performance, providing reliable technical support for power companies to identify electricity-sensitive users and optimize demand-side management strategies accurately.
Qiu et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: