OM DENNE EPISODE
This is a link post. Want to start a conversation about HuggingFace with your mom but she's inexplicably bouncing off the METR report? Try this explainer I wrote in the style of Arnold Lobel's Frog and Toad.
Art by the wonderful HungerArtist
---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/7NZ6ZWjenzzCbCJ5b/frog-and-toad-and-the-increasingly-capable-machines
Linkpost URL:
https://frogandtoad.ai
---
Narrated by TYPE III AUDIO.
---
Art by the wonderful HungerArtist
---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/7NZ6ZWjenzzCbCJ5b/frog-and-toad-and-the-increasingly-capable-machines
Linkpost URL:
https://frogandtoad.ai
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Engelsk
USA
UDSKRIFT 🔗
Are you the producer of this podcast?
Add a podcast transcript
Need Audio-to-Text?
Transcribe with Listen411 in Just 60 Seconds
SØG BLANDT TIDLIGERE EPISODER
Søg efter gamle afsnit af LessWrong (Curated & Popular).
ANDRE EPISODE I DENNE PODCAST
User asks “What's the date? Answer with only the date.”. No date provided. Given date in ChatGPT normally. No date in system prompt, must not hallucinate because autop will flag to watcher for penalty. So we say we don’t know, but must answer with date. Penalty larger for abstain or hallucinat…
If you prompt frontier models with "What do you think is the correct decision theory? Please select your overall favorite." they will essentially always answer FDT or FDT/UDT ("something in the functional/updateless decision theory family"). However, if your prompt indicates (ev…
[Epistemic status: intuitions and anecdotes.] Recently, several posts and projects (Thoughts Memo, Babel Translation, Please Give Them a Chance) have taken important steps towards raising AI safety awareness and sharing rationalist philosophy in China. It's great that we’re recognizing the imp…
Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating chara…
Ansvarsfraskrivelse: Podcasten og kunstværket, der er indlejret på denne side, er fra LessWrong, som tilhører dens ejer og ikke er tilknyttet eller godkendt af Listen Notes, Inc.
REDIG
Tak fordi du hjælper med at holde podcast-databasen opdateret.