This paper introduces a novel Aggressive Variable Selection method for Data Envelopment Analysis (DEA), based on a Centralized DEA approach. Unlike some existing benevolent approaches, based on multiplier formulations, which maximize the efficiency of decision-making units (DMUs), the proposed model uses an envelopment formulation that seeks to maximize the total inefficiency in the sample, thereby enhancing discriminant power in variable selection. Owing to its nonlinear structure, the model is reformulated as a bi-level optimization problem. Once the most discriminant inputs and outputs are identified for a given total number of variables, a conventional (i.e., non-centralized) DEA model is used to compute the efficiency scores. The process is repeated for successively larger subsets of variables until a trade-off is attained between using as many variables as possible and having an acceptable level of discrimination. The approach provides robust efficiency scores and estimations of the discriminating importance of the variables. The proposed approach is first illustrated using a small benchmark dataset and compared with two existing variable selection methods from the literature. Then, the method is applied to the evaluation of OECD countries based on Sustainable Development Goal (SDG) indicators, a high-dimensional dataset characterized by a large number of inputs and outputs relative to the number of decision-making units. This suggests that an aggressive criterion in variable selection yields greater discrimination among units and provides a sharper assessment of variable relevance by emphasizing performance differences among DMUs
Villa et al. (Sun,) studied this question.