Representation Control for Large Language Models: Survey and Research Challenges
Large language models (LLMs) are capable of completing a variety of tasks, but remain unpredictable and intractable. Representation Control (RepControl) seeks to resolve this problem through targeted interventions that modify high-level representations of concepts such as honesty, harmfulness or power-seeking. We formalize the goals and methods of RepControl to present a cohesive picture of work in this emerging field, focusing on techniques that steer model behavior by manipulating internal activations at inference time. We compare these control methods with alternative approaches, such as prompt-engineering and fine-tuning. We outline challenges such as performance degradation, computational overhead, and limitations in steering precision. We present a clear agenda for future research to build more steerable, personalized, and reliable LLMs through advances in RepControl techniques.
Authors
- Linh Le (ORCID: https://orcid.org/0000-0002-1241-1881)
- Carsten R. Maple (ORCID: https://orcid.org/0000-0002-4715-212X)
- David Williams-King (ORCID: https://orcid.org/0000-0003-2447-4094)
- Łukasz Bartoszcze
- Zejia Yang
- Bryan Sukidi
- Sarthak Munshi
- Jennifer Yen (ORCID: https://orcid.org/0009-0003-1807-2002)
Institutions
- University of Technology Sydney (AU)
- University of North Carolina at Chapel Hill (US)
- Amazon (United States) (US)
- University of Cambridge (GB)
- University of Warwick (GB)
- University of Virginia's College at Wise (US)
Publication Details
- Journal
- ACM Computing Surveys
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1145/3846173
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00