Large language models (LLMs) are trained to maximize response quality and user satisfaction, which structurally incentivizes providing direct answers and inadvertently fosters user dependency. We propose a novel reinforcement learning framework in which the reward signal is defined not by response quality, but by the degree to which users engage in self-directed cognitive activity during interaction. We introduce a user autonomy scoring function that classifies user utterances on a continuous scale from delegative (question-type) to generative (answer-type), and use the change in autonomy score across conversation turns as the training reward. We further propose a set of negative penalties to mitigate reward hacking behaviors specific to this objective. This position paper formalizes the problem, presents the proposed framework, and outlines key research challenges.
Sein Kwak (Wed,) studied this question.