ABOUT THIS EPISODE

English
United States

TRANSCRIPT 🔗

SEARCH PAST EPISODES

Search past episodes of LessWrong (Curated & Popular).

OTHER EPISODES IN THIS PODCAST

This is a link post. Want to start a conversation about HuggingFace with your mom but she's inexplicably bouncing off the METR report? Try this explainer I wrote in the style of Arnold Lobel's Frog and Toad. Art by the wonderful HungerArtist --- First published: Sept…
[Epistemic status: intuitions and anecdotes.] Recently, several posts and projects (Thoughts Memo, Babel Translation, Please Give Them a Chance) have taken important steps towards raising AI safety awareness and sharing rationalist philosophy in China. It's great that we’re recognizing the imp…
User asks “What's the date? Answer with only the date.”. No date provided. Given date in ChatGPT normally. No date in system prompt, must not hallucinate because autop will flag to watcher for penalty. So we say we don’t know, but must answer with date. Penalty larger for abstain or hallucinat…
Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating chara…
If you prompt frontier models with "What do you think is the correct decision theory? Please select your overall favorite." they will essentially always answer FDT or FDT/UDT ("something in the functional/updateless decision theory family"). However, if your prompt indicates (ev…
Disclaimer: The podcast and artwork embedded on this page are from LessWrong, which is the property of its owner and not affiliated with or endorsed by Listen Notes, Inc.