ABOUT THIS EPISODE

English
United States

SEARCH PAST EPISODES

Search past episodes of Deep Papers.

OTHER EPISODES IN THIS PODCAST

We discuss Accurate KV Cache Quantization with Outlier Tokens Tracing, a deep dive into improving the efficiency of LLM inference. The authors enhance KV Cache quantization, a technique for reducing memory and compute costs during inference, by introducing a method to identify and exclude outlier t…
What if your LLM could think ahead—preparing answers before questions are even asked? In this week's paper read, we dive into a groundbreaking new paper from researchers at Letta, introducing sleep-time compute: a novel technique that lets models do their heavy lifting offline, well before the…
In this week's episode, we talk about Elastic Reasoning, a novel framework designed to enhance the efficiency and scalability of large reasoning models by explicitly separating the reasoning process into two distinct phases: thinking and solution.  This separation allows for independent alloca…
This week we discuss The Illusion of Thinking, a new paper from researchers at Apple that challenges today’s evaluation methods and introduces a new benchmark: synthetic puzzles with controllable complexity and clean logic.  Their findings? Large Reasoning Models (LRMs) show surprising failure mode…
Disclaimer: The podcast and artwork embedded on this page are from Arize AI, which is the property of its owner and not affiliated with or endorsed by Listen Notes, Inc.