Lilian Weng· @lilianweng · X·· 2024-12-02精选
AI 导读
🦃 感恩节假期结束时,我终于完成了关于 reward hacking 的文章。不好写啊,呼。
推荐理由
OpenAI研究员深度解析奖励作弊机制,揭示LLM自主化部署的关键安全障碍
正文 · 原文
🦃 At the end of Thanksgiving holidays, I finally finished the piece on reward hacking. Not an easy one to write, phew.
Reward hacking occurs when an RL agent exploits flaws in the reward function or env to maximize rewards without learning the intended behavior. This is imo a major blockers for real-world deployment of more autonomous use cases of AI models.
Also would like to call out more research on mitigation strategies for reward hacking, especially in the context of LLMs and RLHF.
👉https://lilianweng.github.io/posts/2024-11-28-reward-hacking/
来源:Lilian Weng · x.com