From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning

Chen, Wei; Huang, Zhen; Xie, Liang; Lin, Binbin; Li, Houqiang; Lu, Le; Tian, Xinmei; Cai, Deng; Zhang, Yonggang; Wang, Wenxiao; Shen, Xu; Ye, Jieping

From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning

Wei Chen, Zhen Huang, Liang Xie, Binbin Lin, Houqiang Li, Le Lu, Xinmei Tian, Deng Cai, Yonggang Zhang, Wenxiao Wang, Xu Shen, Jieping Ye

ICML 2024 pp. 6950-6972

/icml/2024/chen2024icml-yesmen/

Abstract

Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propose to employ supervised fine-tuning (SFT) to mitigate the sycophancy issue, while it typically leads to the degeneration of LLMs’ general capability. To address the challenge, we propose a novel supervised pinpoint tuning (SPT), where the region-of-interest modules are tuned for a given objective. Specifically, SPT first reveals and verifies a small percentage ($<$5%) of the basic modules, which significantly affect a particular behavior of LLMs. i.e., sycophancy. Subsequently, SPT merely fine-tunes these identified modules while freezing the rest. To verify the effectiveness of the proposed SPT, we conduct comprehensive experiments, demonstrating that SPT significantly mitigates the sycophancy issue of LLMs (even better than SFT). Moreover, SPT introduces limited or even no side effects on the general capability of LLMs. Our results shed light on how to precisely, effectively, and efficiently explain and improve the targeted ability of LLMs.

PDF ICML OpenReview Semantic Scholar

Cite

Text

Chen et al. "From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning." International Conference on Machine Learning, 2024.

Markdown

[Chen et al. "From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning." International Conference on Machine Learning, 2024.](https://mlanthology.org/icml/2024/chen2024icml-yesmen/)

BibTeX

@inproceedings{chen2024icml-yesmen,
  title     = {{From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning}},
  author    = {Chen, Wei and Huang, Zhen and Xie, Liang and Lin, Binbin and Li, Houqiang and Lu, Le and Tian, Xinmei and Cai, Deng and Zhang, Yonggang and Wang, Wenxiao and Shen, Xu and Ye, Jieping},
  booktitle = {International Conference on Machine Learning},
  year      = {2024},
  pages     = {6950-6972},
  volume    = {235},
  url       = {https://mlanthology.org/icml/2024/chen2024icml-yesmen/}
}