Kim Isenberg: ⁓ Yeah. We were just talking actually outside with one of the the folks that we know from one of our early access programs who had built a mobile app powered by Gemma three years ago. This would have been impossible, a profitable business. The models can only do ninety-five percent of the things I want them to do in the day to day. And it might look correct. Many of the people are coming from outside of this world of software engineering and they don't know how to kind of double click and make sure that all of those connections that they asked for are actually robust. I mean if you want like production grade systems you need to be careful with some of these things. We want to build the best models we can build. Right. And in that sense our benchmark is ourselves rather than the competitors. We know what the community wants. If you do give everyone the ability to create any sort of app for zero dollars, super fast, completely on device from a mobile phone, like there's there's really no limit to the potential of this. Globally, like not just in the United States, not just in Europe, but everywhere. Hi, I'm Kim Eisenberg, super intelligence editor in chief. And today I'm joined by Omar and Paige from Google DeepMind. Omar leads developer experience. He's the person making sure models like Gemini and Gemma actually reach developers. And before DeepMind, he ran platform and community at Hugging Face. And Paige is the technical, you unit technical lead for creator and developer experiences, driving the product strategy behind Gemini. And she previously worked at GitHub. Co-pilot, right? Omar Paige, great to have you both here. Great to be here. Thank you so much for coming to I.O. this year. Yeah, thank you for having us here. Thank you so much for inviting me. So it's been big days of announcements. Before we get into specifics, what's the thing you should showed at IO that you're personally most excited about? So I I feel like the the experiences that we just added in AI Studio are pretty compelling. ⁓ we talked a little bit earlier about managed agents, which our our colleagues have been ⁓ sort of shipping out. ⁓ they give you the ability to have ⁓ these agents that are running 247 on managed cloud VMs, which is pretty magical. And you can invoke them just by asking in natural language for things that you would like to have created. And then also I'm really jazzed about ⁓ being able to vibe code Android apps in AI Studio. Yeah, for me, something that has been quite exciting to see over the last few months is that the models are finally there. Like the models can finally do 95% of the things I want them to do in the day to day. The next layer there is the agency hardness, right? Which is pretty much what Patch is talking about. The ⁓ in this case, anti gravity as a identity hardness, not just as the ⁓ coding ⁓ platform, but really as a agency harness that is enabling all of these different experiences from the managed agents in AI Studio to buy coding in AI Studio to a bunch of different products across Google. And I think that's very, very exciting. And it's I love that you called out that it's like a singular agent harness. Whereas in the before times, I think there were many different ones that you could choose from. And now everything from Gemini Spark to AI Studio to Anti-Gravity is running the same tool in the Right, makes totally sense. ⁓ Omar, you spent years at Hugging Face, basically the home of open source AI, right? Now you're at Google DeepMind building the models that get distributed through platforms like Hugging Face. How did that switch change how you think about what developers actually need? Tough question, right? I mean, it's ⁓ it's an interesting moment in the industry because ⁓ so I I do have a bias to open models and I've been working in the open models space for many, many years. ⁓ going back to my previous answer as well is the models can finally do the things, right? ⁓ in the context of Gemini, Gemini can do 95% of the things I want it to do in my day-to-day. The context of models such as Gemma, they are on device, much smaller, very different use cases, right? Like privacy heavy ⁓ areas, but these are very powerful models. Something that has been very exciting ⁓ to work at Google is how the models get distributed, not just by how they are in ⁓ in Hogging Face, but also how they are integrated into Chrome, into Android. So for example, Yemi 9 Nano is built on top of Gemma. And of course, even like ⁓ via ⁓ AI Studio, right? Like you can build with Gemma. So yeah, I think that part is quite exciting. ⁓ Google is in a very good place to distribute those models to millions of people, ⁓ directly embedded in their experiences that they use in their day to day, which is very a very different experience than going to Hogin Face maybe to download and fine-tune the model. You already brought up Gemma, so let's ⁓ stick with Gemma for a moment, right? Gemma now ships under Apache 2 license, which is genuinely a permissive open source license. ⁓ that's a significant significant decision. Why did Google make the choice and what does it mean for developers in practice? Yeah, so we have been doing things ⁓ in a very community centric way for quite some time. And this is not just Gemma, right? Like Paige can also share about Jeminai. ⁓ we have been talking with developers, talking with startups, and we keep hearing this kind of feedback, right? ⁓ what are the capabilities that they want, what's the license that they want, which are the things that they want, and we're iterating quickly, we're working as a startup. incorporating that feedback into the roadmap and really reacting based on that. Yeah. I agree. I I think that the Gemma early access program has been a really strong channel and conduit for getting feedback from the community, either people who are building with Gemma in the academic world or in the sciences world or ⁓ you know, partners like Unsloth or Olama. Yeah. ⁓ and Apache too is such a a nice gift to the community to be able to give them the confidence that they can build their businesses around Gemma. ⁓ It's really really powerful to see. Yeah. A a big part of Gemma is building a foundation for the ecosystem to build on top of, right? And having a license they understand that they feel comfortable with is critical for sovereign use cases, for AI for good, for a bunch of day-to-day use cases as well. We were just talking actually outside ⁓ with one of the the folks that we know from one of our early access programs who had built a mobile app powered by Gemma ⁓ that could do kind of tool calling and retrieval. ⁓ all completely on his device. And we were discussing like, you know, three years ago this would have been a profitable business well, impossible, a profitable boot business, you know. ⁓ and he was even ⁓ you know, thinking about commercializing it. So and now it's great to know that based on these permissive licenses he would have the option to do them. Exactly. ⁓ I mean this is a good ⁓ point to talk about because the landscape is very competitive right now, right? Gemma, Quinn, Mistral and more If I'm the if I'm a developer starting a project today, ⁓ why should I pick Gemma? What's the honest case here? I mean, I can say nice things about Gemma. And I I feel like I'm slightly less biased because like Omar Omar's part of the the Gemma team and I'm just like a Gemma super fan. Sorry, so part of the Gemma thing. I'm also a super fan, actually. Well, it's the so Gemma it has ⁓ many of the same capabilities that are magical about the Gemini model family. So it's able to understand video now, audio, images, multilingual, multilingual. It's really, really good at tool calling. Yep. ⁓ it comes in sizes that make sense if you're trying to productionize AI. So the two B, the four B, the ⁓ the two larger sizes, one of them is even a dense model, one of them is a mixture of experts model. ⁓ and it's been really, really nice to just kind of see the creative ways that people in the community are experimenting with it. It also has a really nice ⁓ kind of context window to 156k. Yeah. And then also it has ⁓ also it has like this ⁓ support for over 140 different languages and even more. ⁓ it's just that they haven't been kind of officially tested or kind of blessed. ⁓ yeah, like definitely like, you know, you might get it to be able to say some Klingon And Slay and or some Elvish and ⁓ and you know who knows. So yeah, in the development of Gemma for us is what is important is to make models that are not just the best models out there, but models that are actually usable by the community, right? That's why we care about developer friendly sizes. So we care a lot about building the most capabilities or intelligence per parameter, the most intelligent model per watt, ⁓ and making the models in such a size that they can run in devices that people have in their day to day. So the smallest model models I have here, my pixel phone from years ago, like I can be run the model here, right? Or ⁓ in a laptop or in the co gaming computers. Yeah. Not you you should not need to have like a huge ⁓ number of GPUs to be able to run some of these models at home. And sh and Shinova from ⁓ from the Hugging Face team just recently got the two billion perimeter version or was it four billion perimeter version? Four, I think. The four billion perimeter version running with Transformer JS in the browser controlling a robot. Which is wild. Yeah, yeah. Yeah, you can do things like a Raspberry Pi. You can put the model in extremely ⁓ hardware constrained setups, ⁓ and just a model and it's not a slow ⁓ experience. Like it can actually work quite I mean, yeah, but this is precisely ⁓ my my next question. ⁓ it's designed to run on a single consumer GPU exact bytes, but it also can run ⁓ on a phone, for example. ⁓ could you explain Why the size range matters and what the difference is between running a model in the cloud versus running it yeah on your own device. Okay, where is the advantage here? Yeah, so there are three different ⁓ architectures or sizes. So the E2B and E4B. So those are the two smallest models. They were architecturally designed to be very efficient to run in a phone. ⁓ so they use an architecture called PLE per layer embedding, which is specifically done to be able to run very efficiently. in in the phones. So and that was done in partnerships with different OEMs and and so on. And then Paige mentioned before, like there's both a dense and a MOE. So dense model is like the 30, the 31B. It's the most ⁓ raw intelligence is the largest model. If you want the best open model from us, ⁓ you would go and use that one. But for many deployment setups you want very, very fast in friends. And that's where a mixture of experts would come in. So the MOE is a 26 B ⁓ parameter model, so it's quite large. ⁓ But it only four billion parameters actually activate. And that means that you have almost the same intelligence as a twenty seven B model, but only four billion like you get the speed of a four B model. I'm simplifying a bit things, but that's more or less how it works. So ⁓ we're trying to serve different kind of age use cases, right? From the people that want to run the models in their in their pockets to people that want to deploy in production setups to people that want to Build the most ⁓ powerful model without worrying too much about the latency. Yeah. And also the energy consumption. Like the the smaller models, if you're if you're operating on a cell phone, you want something that doesn't drain your battery power really quickly. ⁓ and then also from an energy efficiency perspective, I know people are starting to care really deeply about that. ⁓ as well as ⁓ a I think distributed inference can still be a challenge for many folks. And so being able to fit on a single commodity GPU is is a lot easier to architect than than multiple GPUs. The electricity consumption, I think is super important. ⁓ because in a phone setup, like you don't want the battery to exactly to be drained, right? So one thing that ⁓ we don't talk that much about that probably we should talk more about is the the how how many tokens a Yama generates, right? Because one part is okay the model is very intelligent, but another question which is very important in on device setups is How many tokens the model needs to generate to get to the right answer. Right. And that's something that we we benchmark, we validate, we train specifically for to make sure that the density of a reasoning that it can do when it's doing reasoning is very strong. So the model can do very good reasoning without generating tens of thousands of tokens. Yeah. And you can see that in the efficiency of the tool calls. Like I I really hope that Andrew opens sources or at least like shares his afflict the world so that ⁓ but it was ⁓ it was relatively speedy in terms of the the tool calls needed to invoke different ⁓ different applications. ⁓ and people can also test out the Google AI Edge Gallery app. Yeah, so you g ⁓ so if you use iPhone or use Android, there is the AI Edge Gallery, it's in both Play Stores. So you can go download it mod ⁓ the app ⁓ and you can use it in on an airplane with this guy. Very cool. I I love it. Having like a model running locally ⁓ Like to say when I'm on a plane, right, having a model, it it's perfect. It's amazing. First thing the first thing I did downloading it. It feels magical. One of one of our colleagues, Olivier, was also showing a demo yesterday of ⁓ being able to have kind of the video input feed for a phone working on a hike. So ⁓ like you can ask about signs, even signs in different languages, like what they mean, which direction to go. I all these things. Yeah, that it to me feels magic. It's it's like the features here, right? Yeah, exactly. But ⁓ to continue, ⁓ are we moving toward a world where instead of one big general purpose model, developers build with a toolkit of smaller, specialized models, is it the future you're designing for? So more specialized model instead of general purpose models. I would ⁓ so so I really love the this sort of behavior that we see right now with agents that are figuring out what needs to be the model. generating a whole bunch of tokens versus what should be a tool call. ⁓ like what tools, what ⁓ sort of options do I have available that doesn't require me to generate a lot of tokens and run through all of this reasoning by myself. ⁓ because I I think that leads to a much more kind of efficient AI future if you if you do have this space where ⁓ you know you have a powerful reasoning model like Gemini, it's able to break down a complex task into you know, 50 different steps. And then it's able to decide like, all right, 30 of these steps I can use Gemma, 20 of them or 10 of them I can use these tools that aren't even like relying on AI at all. And then these other 10 might need to be calls to Gemini. ⁓ but that dramatically reduces, you know, your energy consumption, the time that it might take to achieve an outcome, the cost profile that you might have for a problem. ⁓ and I I really hope that that's the future. ⁓ We're we're already starting to see it a little bit. But the ⁓ but I I think that ⁓ DeepMind is also really great at kind of building these models that are capable of doing so many things. Yeah, that's a very interesting part, right? Like ⁓ how much do you think we'll go in a more unified direction? And Omni is a great move in that direction. Like Gemini Gemini is pretty powerful in the sense that it can understand video and images and audio and text and code and it can also output multiple modalities. So I can output audio, output images, output text, output code, and now video for like this new Omni exploration that we're doing. ⁓ so I I think that, you know, there's space for both. ⁓ but from a from a person who has probably too many hobby projects. Like it's and a very large personal cloud bill. Like I I really love this idea that we could ⁓ like dramatically reduce Page's cloud bill. And like use use a constellation of models as opposed to just one singular. Yeah. Yeah. Yeah. For Yemma, it's been interesting to see. We with Yemma four the mo the the base model is just so strong that we see less of a need for people to fine tune the model. So with Gemma three, we had tons of partners that were doing fine tuning on top of it. With Yemma4, like people were getting state of the art, ⁓ results ⁓ for open models. In their benchmarks without any additional fine tuning. And it was the same in the early days of the Gemini model or the Gemini models as well. Like the like at the very beginning they had to have ⁓ so people would fine tune Gemini to create Medpalm or ⁓ to create kind of ⁓ like sp specific to coding versions of Gemini. And now like the the benchmarks for those those tasks are just exceeded by base Gemini. I'm just curious. ⁓ how was the feedback so far from the community? And are there some wishes? ⁓ they were like, okay, this this needs to be like a developed further in Jammer or ⁓ and from the community? Yeah, so Sundar announced I think twenty million downloads the week after the launch. ⁓ by now we have over one hundred million downloads in six weeks. So the community reception has been very positive and With GMA3, when we release M3, we had lots of feedback from the commit. Like people were telling us function tooling doesn't work well, system instruction doesn't work well, all of these things does not work well. With GMA4, we really ⁓ incorporated all of that. So in general, the reception has been extremely positive. We know on the agentic side of things, we still need to we we still want to keep pushing the frontiers, but always constraining within the the same size, right? I think that's something I'm very excited about. If you compare similarly sized Gemma models to Gemma four, so we had Gemma two twenty seven B, Gemma three, twenty seven B, Gemma four, thirty one B, we kept seeing an improvement across all of the LM Arena ⁓ verticals. So we think that we can keep pushing significantly the capa the identity capabilities ⁓ of Gemma without having to keep scaling up the model size. ⁓ And the team is ⁓ one of the things that's pretty magical too is that so many of the the Gemma team members are on social media, they're in the EAPs, they're answering feedback all the time, listening to feedback. Building with open tools. Exactly. Yeah. Yeah, I think that's something very important, right? Like people need to build with ⁓ with what developers are ⁓ building as well, right? ⁓ and I think that's quite critical for the success. ⁓ Gemma cool. Paige. You've built developer tools at GitHub with Copilot and now at Google for AI developers. What's the biggest gap you see right now between what frontier AI models can do and what a typical developer can actually build with them? Interesting. So so I I think the ⁓ one of the the most one of the most interesting gaps that I've seen, and I guess ⁓ for context. I I started ⁓ contributing to open source projects a long time ago. What we were just talking like almost two decades ago. ⁓ first in the scientific computing space, ⁓ like NumPy and SciPy and ScikitLearn, and then later on things like TensorFlow and ⁓ and some of the other machine learning frameworks that we have at Google. ⁓ but it there's always been kind of this gap between like understanding the complexity of a system. ⁓ and ⁓ being able to ⁓ sort of under ⁓ understand like where might the gaps in this system be. ⁓ and I think today one of the the things where I see developers continuously get frustrated, and I would love to hear Omer's perspective on this as well, is that people might ask in a tool like Anti-Gravity or AI Studio or Cursor or whatever it might be, please ⁓ do this task for me. And it might require connecting to a database or a system. or ⁓ you know creating a specific kind of database, ⁓ incorporating something like OAuth ⁓ or pulling in data via search or or some other tool call. and the the IDE or the tool will give a response back and it might look correct. ⁓ but many of the people are coming from outside of this world of software engineering and they don't know how to kind of double click and make sure that all of those connections that they asked for actually robust. ⁓ that it is a connection to a database as opposed to just synthetic data that was created behind the scenes. ⁓ it is actually doing, you know, a connection to workspace ⁓ as opposed to just kind of like hallucinating some or fabricating some of the the workspace ⁓ sort of examples. And I think we're we're working really hard in AI Studio to make sure that those connections exist and that they're ⁓ invoked explicitly. ⁓ but it's really, really hard for earlier career devs to ⁓ to be able to double check that that some of these things actually are correct. So so I think that people are right now expecting the models to do a lot and those expectations keep increasing over time. Yeah. and the models really can do a lot. It's just that right now we're in this strange interstitial state where we also need to have ⁓ you know, a healthy amount of skepticism that it does all of the things that you ask for. Yeah. ⁓ and and that's that's a gap that I that I see people continuously getting getting snagged by that hopefully we'll get better with time. And we're trying to close an AI studio. Yeah. Yeah, it's an interesting point, right? Because we're at a stage in which it's never been easier to be with AI. And at the same time, if you want like production grade systems, you need to be careful with some of these things. And I think ⁓ as you mentioned, like non technical people or tech adjacent people that are starting to build with AI, but also students, right? Like what is happening now with computer science students, once they are are out there, how are they going to build, how their development workflows will look like, ⁓ will be something very interesting to To follow and to be the right tooling for them as well, right? And to make sure that they're taken care of. So if they do accidentally, you know, expose an API key or expose password in plain text, that those things are caught for them as opposed to like foot guns that they find out afterwards. Yeah. Yeah. This is actually exactly my follow-up question. So for someone who isn't in machine learning engineer but wants to build something with AI like ⁓ journalism or business owner, designer, whatnot. ⁓ how far can they actually get with Google Ads Studio today without writing code? ⁓ god. Very far. Super far. Like and and like and much further since yesterday. Yeah, yeah, yeah. Absolutely. So so they can have databases in OAuth. We've added workspace support. So you can connect Gmail and Calendar. and things like mobile development. You could deploy to an Android app. You can export to anti-gravity if you need to. ⁓ we've also got grounding with Google search, grounding with Google Maps. ⁓ URL context, which is effectively like retrieval for free. So you don't have to stand up a vector database. Yeah. ⁓ and then you have the I we we didn't launch this yesterday, but we launched this in the last two months. You have this ⁓ editor and annotator mode in which you can just find more things around. And it's learning all of these multimodal capabilities of Gemini to keep iterating on the product without writing a single line of code. And also design features. So it's like ⁓ once you once you ⁓ describe an app that you want to create, like it gives you five different design options. So like if you're like me and you can't design and front end at all, like it it gives you, you know, ⁓ the ability to say this one looks beautiful and then and then use it in your app. I I love the answer. I love it, I love it. So there's ⁓ as you know, a global race in the open. models, right? Chinese labs, European companies, Ms. R, for example, and now Google are already already using very competitive open modes. How does the DeepMind team think about this competition? ⁓ is it a race you need to win or does more competition help everyone? The way I say it is that we want to build the best models we can build, right? And in that sense our benchmark is ourselves rather than the competitors. We know what the community wants. ⁓ And we're iterating based on that. So in the context of Gemma, so in the context of open models, again we want to build the strongest models that are developer friendly, consumer friendly, that can run in consumer devices. In the context of Gemini and the OPNI, Bio, Liria, and Nano Banana, like the whole Gemini family. Again, the goal is to build the best models we can build ⁓ based on all of the feedback we get from the community. Yeah. And I and I do feel like the The world will be a better place, the more open models that exist. Couldn't agree more. So it's it's kind of exciting to see how all of these, ⁓ all of these are are getting deployed out into the world in fine-tuned species. Okay, so if we sit here at Google IO next year, what will have changed the most about how people build with AI? Tough question again. It's a fun one. It's a fun one. I mean ⁓ The models are getting very, very good. The Agenti harnesses are also getting to a very, very good place. What I'm the most excited for the next 10 months, six months is to see more and more of these non-traditional developer audiences starting to build with AI. And I think that will help us shape how we want the roadmap to look like, right? So I don't know what we're going to ship in a year from now, but I'm sure like seeing how people that don't come from a developer background build with AI. will help inform ⁓ this significant there. I love that. And I I was going to say something really similar at late. We just announced an AI Studio mobile app at I.O. ⁓ this year, ⁓ which is, you know, massive, massive market. Like there are many people who have cell phones or who ⁓ kind of ⁓ are very devoted to their cell phones but don't necessarily have a laptop at home. ⁓ and given how much ⁓ there is still left to build, how much creativity and passion that folks have around the world, especially out of the engineering space. ⁓ I can't wait to see what they create. ⁓ and also like see more people rely on on device models or working with anti-gravity or with mobile devices. I think that would be really interesting to be interesting for sure. Yeah. Because if you if you do give everyone the ability to create any sort of app for zero dollars super fast. completely undevied from a mobile phone. ⁓ like there's there's really no limit to the potential of this ⁓ globally, like not just in the United States, not just in Europe, but everywhere. Yeah. So cool. Yeah. Omar Paige, thank you so much for the conversation. I really, really appreciate it. It had much fun. Thank you so much so much for having us. Thank you so much.
درباره این اپیزود
In this exclusive interview from Google I/O, I speak with Omar Sanseviero and Paige Bailey from Google DeepMind about the rapidly evolving AI landscape.
We discuss the rise of local models, the growing importance of open source and open models, the role of developer communities, and how global competition — especially from China — is shaping the next phase of artificial intelligence.
A conversation about where AI is heading next: from frontier labs to local inference, from closed systems to open ecosystems, and from model releases to real-world developer adoption.
انگلیسی
ایالات متحده آمریکا
رونوشت 🔗
جستجوی اپیزودهای گذشته
اپیزودهای قبلی The Superintelligence Podcast را جستجو کن.
سلب مسئولیت: پادکست و آثار هنری تعبیه شده در این صفحه متعلق به Kim Isenberg & Peter Thum است که متعلق به صاحب آن است و به Listen Notes، Inc وابسته یا تایید نشده است.
ویرایش
از کمک شما برای بروز نگهداشتن پایگاهدادههای پادکست سپاسگزاریم