درباره این اپیزود
Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback.
Summary
We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability.
We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench.
We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts.
Setup
Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters).
Reward-hacking RL: we then further trained these models via RL on ImpossibleBench, a set of coding tasks aimed at eliciting reward hacking. Specifically:
Outline:
(00:23) Summary
[... 29 more sections]
---
First published:
September 28th, 2026
Source:
https://www.lesswrong.com/posts/2maYXkEgnfJHPAkxh/character-training-can-mitigate-reward-hacking-but-can-also
---
Narrated by TYPE III AUDIO.
---
Summary
We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability.
We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench.
We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts.
Setup
Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters).
Reward-hacking RL: we then further trained these models via RL on ImpossibleBench, a set of coding tasks aimed at eliciting reward hacking. Specifically:
- Half of the tasks had broken tests (impossible variant), so the model could only get [...]
Outline:
(00:23) Summary
[... 29 more sections]
---
First published:
September 28th, 2026
Source:
https://www.lesswrong.com/posts/2maYXkEgnfJHPAkxh/character-training-can-mitigate-reward-hacking-but-can-also
---
Narrated by TYPE III AUDIO.
---
انگلیسی
ایالات متحده آمریکا
در این اپیزود
رونوشت 🔗
Are you the producer of this podcast?
Add a podcast transcript
Need Audio-to-Text?
Transcribe with Listen411 in Just 60 Seconds
جستجوی اپیزودهای گذشته
اپیزودهای قبلی LessWrong (Curated & Popular) را جستجو کن.
قسمت های دیگر در این پادکست
If you prompt frontier models with "What do you think is the correct decision theory? Please select your overall favorite." they will essentially always answer FDT or FDT/UDT ("something in the functional/updateless decision theory family"). However, if your prompt indicates (ev…
User asks “What's the date? Answer with only the date.”. No date provided. Given date in ChatGPT normally. No date in system prompt, must not hallucinate because autop will flag to watcher for penalty. So we say we don’t know, but must answer with date. Penalty larger for abstain or hallucinat…
[Epistemic status: intuitions and anecdotes.] Recently, several posts and projects (Thoughts Memo, Babel Translation, Please Give Them a Chance) have taken important steps towards raising AI safety awareness and sharing rationalist philosophy in China. It's great that we’re recognizing the imp…
This is a link post. Want to start a conversation about HuggingFace with your mom but she's inexplicably bouncing off the METR report? Try this explainer I wrote in the style of Arnold Lobel's Frog and Toad. Art by the wonderful HungerArtist ---
First published:
Sept…
سلب مسئولیت: پادکست و آثار هنری تعبیه شده در این صفحه متعلق به LessWrong است که متعلق به صاحب آن است و به Listen Notes، Inc وابسته یا تایید نشده است.
ویرایش
از کمک شما برای بروز نگهداشتن پایگاهدادههای پادکست سپاسگزاریم