เกี่ยวกับตอนนี้

ภาษาอังกฤษ
สหรัฐอเมริกา

สำเนาบทสนทนา 🔗

ค้นหาตอนที่ผ่านมา

ค้นหาตอนที่ผ่านมาของ LessWrong (Curated & Popular)

ตอนอื่นๆ ใน PODCAST นี้

Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating chara…
[Epistemic status: intuitions and anecdotes.] Recently, several posts and projects (Thoughts Memo, Babel Translation, Please Give Them a Chance) have taken important steps towards raising AI safety awareness and sharing rationalist philosophy in China. It's great that we’re recognizing the imp…
User asks “What's the date? Answer with only the date.”. No date provided. Given date in ChatGPT normally. No date in system prompt, must not hallucinate because autop will flag to watcher for penalty. So we say we don’t know, but must answer with date. Penalty larger for abstain or hallucinat…
If you prompt frontier models with "What do you think is the correct decision theory? Please select your overall favorite." they will essentially always answer FDT or FDT/UDT ("something in the functional/updateless decision theory family"). However, if your prompt indicates (ev…
ข้อสงวนสิทธิ์: พอดแคสต์และอาร์ตเวิร์คที่ฝังอยู่ในหน้านี้มาจาก LessWrong ซึ่งเป็นทรัพย์สินของเจ้าของและไม่มีส่วนเกี่ยวข้องหรือรับรองโดย Listen Notes, Inc.