When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models Under review
Do LLMs have desires? Choice paradigms indicate that they express a coherent utility structure, but that doesn't entail that LLMs will act in accord with these reported preferences when given the opportunity. We develop an experimental paradigm composed of a variety of writing tasks and an LLM judge panel that can discern differences in model output quality. We find that LLMs can be motivated to produce outputs of varying quality by 1) exhorting them to try hard, 2) prompting them to play a role of someone adept at the task, or 3) attaching performance to harmful outcomes. However, no LLM, on any task, varies its output quality when attaching its performance to outcomes the choice paradigm indicates it has high utility for.