OM DENNE EPISODE
Many AI features in web apps send the user's audio, images, or text to a server and wait for the result. That adds network latency, moves data off the device, and ties the experience to a connection. Transformers.js gives developers another option: run the model in the browser.
In this episode, Mike Grabowski speaks with Nico Martin, Open Source ML Engineer at Hugging Face, about what browser AI can do today. They break down the Transformers.js pipeline API, ONNX Runtime, WebGPU acceleration, model downloads, browser caching, CPU fallbacks, and performance across devices.
Nico explains when local inference can beat a server round trip, why model size and hardware variation shape the user experience, and how browser code can handle parts of an agent workflow. He also shares early work on structured output and a custom WebGPU inference engine, where current experiments point to 5x to 10x speedups.
For developers deciding which AI tasks belong in the browser and which need the cloud, this episode maps the tradeoffs through speech recognition, background removal, embeddings, local tool calling, and on-device agents.
Check out episode resources on our website ➡️ https://clstk.com/4zMwzq9
Catch more React Universe On Air episodes 🎧 https://clstk.com/4zNCV8W
Sign up for our newsletter ✉️ https://clstk.com/4h2b8Ks
Follow us on X 🐦 https://x.com/callstackio
Chapters:
0:00 Welcome to React Universe On Air
1:04 Nico Martin's work at Hugging Face
2:04 What Transformers.js is
3:30 Browser AI beyond large language models
6:00 Bringing Python ideas into a JavaScript API
10:08 How the Transformers.js pipeline API works
14:23 Choosing compatible models and architectures
15:46 Why Transformers.js uses ONNX Runtime
19:01 CPU inference, WebGPU, and browser coverage
22:57 Model downloads, browser caching, and UX
28:28 Startup time and performance across devices
30:57 Why Transformers.js focuses on the browser
34:07 What developers misunderstand about browser AI
38:44 Building AI agents in the browser
43:19 Local fallbacks and hybrid AI
44:28 The browser AI roadmap
47:04 Hugging Face's custom WebGPU inference engine
50:57 Why speedups matter on constrained hardware
52:07 Nico's favorite browser AI use cases
In this episode, Mike Grabowski speaks with Nico Martin, Open Source ML Engineer at Hugging Face, about what browser AI can do today. They break down the Transformers.js pipeline API, ONNX Runtime, WebGPU acceleration, model downloads, browser caching, CPU fallbacks, and performance across devices.
Nico explains when local inference can beat a server round trip, why model size and hardware variation shape the user experience, and how browser code can handle parts of an agent workflow. He also shares early work on structured output and a custom WebGPU inference engine, where current experiments point to 5x to 10x speedups.
For developers deciding which AI tasks belong in the browser and which need the cloud, this episode maps the tradeoffs through speech recognition, background removal, embeddings, local tool calling, and on-device agents.
Check out episode resources on our website ➡️ https://clstk.com/4zMwzq9
Catch more React Universe On Air episodes 🎧 https://clstk.com/4zNCV8W
Sign up for our newsletter ✉️ https://clstk.com/4h2b8Ks
Follow us on X 🐦 https://x.com/callstackio
Chapters:
0:00 Welcome to React Universe On Air
1:04 Nico Martin's work at Hugging Face
2:04 What Transformers.js is
3:30 Browser AI beyond large language models
6:00 Bringing Python ideas into a JavaScript API
10:08 How the Transformers.js pipeline API works
14:23 Choosing compatible models and architectures
15:46 Why Transformers.js uses ONNX Runtime
19:01 CPU inference, WebGPU, and browser coverage
22:57 Model downloads, browser caching, and UX
28:28 Startup time and performance across devices
30:57 Why Transformers.js focuses on the browser
34:07 What developers misunderstand about browser AI
38:44 Building AI agents in the browser
43:19 Local fallbacks and hybrid AI
44:28 The browser AI roadmap
47:04 Hugging Face's custom WebGPU inference engine
50:57 Why speedups matter on constrained hardware
52:07 Nico's favorite browser AI use cases
Engelsk
USA
UDSKRIFT 🔗
Are you the producer of this podcast?
Add a podcast transcript
Need Audio-to-Text?
Transcribe with Listen411 in Just 60 Seconds
SØG BLANDT TIDLIGERE EPISODER
Søg efter gamle afsnit af React Universe On Air.
ANDRE EPISODE I DENNE PODCAST
Should you be afraid of monorepos? Not with Nx. Tune in and learn how to scale apps without scaling pain.
In this episode of React Universe On Air, Łukasz Chludziński chats with Jordan Powell from Nx to explore what it really takes to build and manage a monorepo at scale. From dependency graphs to…
Many React Native apps ship without full observability. The result? Blind spots in performance, crashes, and user behavior once your app is in the wild.
In this episode of React Universe On Air, Łukasz Chludziński sits down with Jonathan Munz (Senior Software Engineer at Embrace) and Adam Horodysk…
It’s a practical, developer‑focused look at the early days of React Native on VR: what works today, what’s different from mobile, and why so many app ideas translate better than you might expect.
Chapters:
00:00 Introduction to React Native on Meta Quest
01:10 Meet the guests: Jan Jaworski and Ram…
The NFL runs on tight deadlines. Every kickoff and Super Bowl demands that millions of fans get flawless live video, stats, and fantasy features, whether they’re on iPhones, Xboxes, or connected TVs.
In this episode of React Universe On Air, Mike Grabowski talks with Michael Blanchard (Director of…
What really happens when React Native’s core contributors gather behind closed doors?
In this episode of React Universe On Air, recorded live at React Universe Conf 2025, host Łukasz Chludziński sits down with Joel Arvidsson, Jay Meistrich, and Thibault Malbranche to unpack the hottest debates and…
Ansvarsfraskrivelse: Podcasten og kunstværket, der er indlejret på denne side, er fra Callstack, som tilhører dens ejer og ikke er tilknyttet eller godkendt af Listen Notes, Inc.
REDIG
Tak fordi du hjælper med at holde podcast-databasen opdateret.