# Jacob Steinhardt

인기 팟 캐스트 에피소드
업데이트 됨:

AF - Approaching Human-Level Forecasting with Language Models by Fred Zhang

Link to original articleWelcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Approaching Human-Level Forecasting with Language Models, published by Fred Zhang on February 29, 2024 on The AI Al…

AF - Mechanistic Interpretability Workshop Happening at ICML 2024! by Neel Nanda

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Mechanistic Interpretability Workshop Happening at ICML 2024!, published by Neel Nanda on May 3, 2024 on The AI Alignment Forum. Announcing…

AF - Mechanistic Interpretability Workshop Happening at ICML 2024! by Neel Nanda

Link to original articleWelcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Mechanistic Interpretability Workshop Happening at ICML 2024!, published by Neel Nanda on May 3, 2024 on The AI Ali…

Ep 14 - Interp, latent robustness, RLHF limitations w/ Stephen Casper (PhD AI researcher, MIT)

We speak with Stephen Casper, or "Cas" as his friends call him. Cas is a PhD student at MIT in the Computer Science (EECS) department, in the Algorithmic Alignment Group advised by Prof Dylan Hadfield-Menell. Formerly, he worked with the Harvard Kreiman Lab and the Center for Human-Compatible AI (C…

EA - Language models surprised us by Ajeya

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Language models surprised us, published by Ajeya on August 30, 2023 on The Effective Altruism Forum.Note: This post was crossposted from Pl…

Do AI As Science Instead

Few AI experiments constitute meaningful tests of hypotheses. As a branch of machine learning research, AI science has concentrated on black box investigation of training time phenomena. The best of this work is has been scientifically excellent. However, the hypotheses tested are mainly irrelevant…

Arxiv Paper - Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation

In this episode, we discuss Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation by Danny Halawi, Alexander Wei, Eric Wallace, Tony T. Wang, Nika Haghtalab, Jacob Steinhardt. The paper highlights security risks in black-box finetuning interfaces for large language models and intro…

"Dario probably doesn’t believe in superintelligence" by RobertM

Epistemic status: I think this is true but don't think this post is a very strong argument for the case, or particularly interesting to read. But I had to get 500 words out! I think the 2013 conversation is interesting reading as a piece of history, separate from the top-level question, and recomme…

The State of Hebrew Type Design

“Try and tell me why design doesn't matter.” 🤖🥊🤨  ​​Franziska (Franzisca) Baruch (1901–1989) became interested in Hebrew letterforms before she even spoke it. A letterer and graphic designer, she is likely best known for her Hebrew typeface Stam (Berthold, 1926); depending on how you count, she mad…