ABOUT THIS EPISODE

English
United States

SEARCH PAST EPISODES

Search past episodes of Deep Papers.

OTHER EPISODES IN THIS PODCAST

We discuss Accurate KV Cache Quantization with Outlier Tokens Tracing, a deep dive into improving the efficiency of LLM inference. The authors enhance KV Cache quantization, a technique for reducing memory and compute costs during inference, by introducing a method to identify and exclude outlier t…
In this week's episode, we talk about Elastic Reasoning, a novel framework designed to enhance the efficiency and scalability of large reasoning models by explicitly separating the reasoning process into two distinct phases: thinking and solution.  This separation allows for independent alloca…
This week we discuss The Illusion of Thinking, a new paper from researchers at Apple that challenges today’s evaluation methods and introduces a new benchmark: synthetic puzzles with controllable complexity and clean logic.  Their findings? Large Reasoning Models (LRMs) show surprising failure mode…
The authors of the new paper *Self-Adapting Language Models (SEAL)* shared a behind-the-scenes look at their work, motivations, results, and future directions. The paper introduces a novel method for enabling large language models (LLMs) to adapt their own weights using self-generated data and trai…
Disclaimer: The podcast and artwork embedded on this page are from Arize AI, which is the property of its owner and not affiliated with or endorsed by Listen Notes, Inc.