Section Insights
Introduction and Goals
What is the purpose of this presentation?
The presentation aims to provide insights into DeepMind's new models, demonstrate their usage in projects, and encourage audience interaction through questions.
- The session will include live demos and Q&A.
- DeepMind focuses on responsible AI development.
- The goal is to make the presentation useful for attendees.
DeepMind's Mission and Recent Developments
What is DeepMind's mission and what recent developments have occurred?
DeepMind's mission is to build AI responsibly for humanity's benefit, with recent releases including new models and features like a computer use API and managed agents.
- DeepMind emphasizes AI for science and responsible usage.
- Recent model releases include Gemini 3.5 Flash and others.
- The presentation will cover various AI applications.
Gemma 4 Overview
What is Gemma 4 and its capabilities?
Gemma 4 is an open model family with various sizes, allowing users to download, fine-tune, and expand its capabilities for diverse applications.
- Gemma 4 includes multiple parameter sizes for flexibility.
- It is Apache 2 licensed, enabling broad usage.
- Community-built applications showcase its potential.
Local Model Execution and Features
How can users interact with Gemma 4 locally?
Users can run Gemma 4 locally in their browsers, allowing for tasks like creating tables and analyzing data without sending information externally.
- Gemma 4 supports local execution, enhancing privacy.
- It can perform various tasks, including language processing.
- The Google AI edge gallery offers a range of skills for users.
Performance and Accessibility of Gemma 4
What are the performance characteristics of Gemma 4?
Gemma 4 models are designed to perform well on commodity hardware, offering cost savings and accessibility for users with lower-end devices.
- Gemma 4 performs above expectations relative to its size.
- It can run on various devices, including mobile and laptops.
- Quantized versions of the model are available for faster performance.
Transcript
0:12 trickle in from other locations and then we'll get started pretty shortly. ideally there will be some time for lots and lots of live demos. so you'll ideally learn some of the things about our new models, how to use them as part of your projects. and then there also should be time for questions, but it also seems like we'll have a small enough audience where if you ask questions throughout the duration of the presentation, we can also do that.
0:38 so end goal is to make this as useful as possible for all of y'all. my name is Paige. I work at DeepMind and excited to show you all what we've been working on too. And as mentioned, we'll get started in just another couple of minutes. It looks like people are still getting scanned at the door.
2:38 All right, let's get started. Cool. So, greetings everyone. As mentioned, my name is Paige. If you have questions throughout the duration of the presentation, please feel free to raise your hand and shout them out. ideally, we'll be able to take some take some as I kind of carine along doing live demos. I also want to give a kind of the requisite caveat at the very beginning that DeepMind's entire mission is to build AI responsibly and for the benefit of humanity. So trying to cast as much light as we can in this in this world that increasingly has increasingly has shade. this translates to things like alpha folds much of the work that we do with our open models including medge gemma and a lot of our robotics and science use cases as well as things like AI for math AI for frontier research. so in general, everything that I show you today, is being used for those projects that are deeply deeply embedded in AI for science. and if any of y'all work in the AI for science space, please feel free to send questions and ask afterwards about how you can use AI to to kind of transform and accelerate that that good work. I don't think it's a secret that Google has been a little bit busy over the course of the last few months. feels like we've been releasing new models, new features, new products every single week. most notably, just recently we released a computer use API which we'll be talking about a little bit.
4:13 something called managed agents which gives you the ability to take kind of a higher order task describe it in a natural language and have a fleet of agents execute on it in a Linux workstation like a sandbox environment where you can add skills you can add kind of dependencies that get pulled in along the way and a whole bunch more. and then also our speechtoech translation API which we'll take a look at in a second. And one of the things that I really really love about Google is that not only are we shipping Frontier models. So you might have heard about Gemini 3.5 Flash which got released at IO. but we've also been focused pretty significantly on our open model releases. so how many folks in the room have heard of Gemma 4? quite a few hands. That's excellent. Gemma 4. Oh yes, absolutely. Go for it.
5:06 Gemma 4 is our open model family. It's the latest iteration. we have many different sizes available. So, there's a two billion parameter, a 4 billion. We just released a 12 billion parameter. and then we also have a 26 billion and a 31 billion mixture of experts and dense model respectively. they're useful for a lot of things and they're also Apache 2 licensed which means you can download them, use them as part of your company, fine-tune them, and kind of, expand on them however you feel like would be most useful. and this Gemma universe is actually pretty cool to see, in terms of people building. I I hate slides, so we're going to see how few slides I can get through today. but this is an example of something that someone from the community has built using Gemma 4 and Fable 5 before it was taken off the market to rewrite some of the kernels. you can load the model directly within the browser. So this is loading Gemma directly in the browser using it via web assembly. It's sandboxed. and then it can do things that that feel pretty magical, right? So this model is this model is currently running completely locally. and if I say something to the effect of create a table with emoji comparing and contrasting all of the Harry Potter books based on which are the funniest and most exciting.
6:44 make sure to give me recommendations on which to read and also maybe incorporate Harry Potter Harry Potter in the methods of rationality or and then you get kind of a a response that feels almost instantaneous. you have kind of the Deathly Hollows, Half Blood Prince, Sword or the Phoenix Goblet of Fire, etc. I'm not sure if I agree with the ranking, but that's but that's a pretty interesting pretty interesting comparison. and then also since it is local to the machine, none of your data is getting sent elsewhere.
7:33 This is not using an API. It's just something that's running locally in the browser with transformers.js and the Gemma 4 model. Gemma 4 is also pretty cool in the sense that you can use Google AI edge gallery in order to analyze some of the model capabilities that we have on device. You can download it and use it with Android and with iOS. and it includes everything from kind of taking images and describing them to automatically transcribing audio in multiple languages. I believe Gemma supports over 140 different languages. and then also testing it out with function calling directly on device. If you haven't had a chance to take a look at the Google AI edge gallery, we have a collection of skills as well. so things like building games, doing haikus, asking queries about weather, being able to schedule events on your calendar for you, that are all available to use just with this AI edge gallery app completely for free and just with locally installed models.
8:37 so if you haven't downloaded it, definitely try it out. There's also a way to use the the kind of accelerator on your local device in order to power the model and do all of the inference work as opposed to just the fall back to the CPU. so if you do have something like a Pixel 10 or a higherend a higherend mobile device, you're already able to use Gemma kind of locally for for all of that work, which is quite cool.
9:07 Gemma 4 just to to give a recap or to place how the model performance versus size shapes up. this is the two of the two kind of largest versions of Gemma that I had mentioned before the 31B and the 26B. they're quite small in comparison to some other models that are on the market but they're performing way above what you might expect. so more than models that are in order of magnitude or larger than their size.
9:37 This is great because it kind of translates into a cost savings perspective. You don't have to worry about distributed inference for local models. you don't have to worry about as large of a GPU footprint. and it's really intended to run super super well on a single commodity GPU as opposed to needing something a little bit more fully featured. You can also run Gemma, some of the variants on even things like Jets and Nanos. and for 12b and below, you can run it locally on your laptop. 2B can even fit handily on mobile devices. We've also released quantized versions. So, the the model that we saw at the very beginning that was running so so blazingly fast was using some of the QAT checkpoints that we have for Gemma. and the QAT checkpoints are quite small when you when you take a look even less than a single gigabyte in size for the the kind of two billion parameter version.
Summary
- DeepMind's mission is to build AI responsibly for the benefit of humanity.
- Recent releases include the computer use API, managed agents, and a speech-to-text translation API.
- Gemma 4 is the latest open model family, available in various sizes (2B, 4B, 12B, 26B, and 31B parameters) and licensed under Apache 2.
- Models can be run locally in a browser using web assembly, ensuring data privacy.
- Gemma 4 supports over 140 languages and can perform tasks like image description and audio transcription.
- The Google AI edge gallery offers a collection of skills for local model use, including game building and event scheduling.
- The models are optimized for performance, requiring less computational power than larger models, making them accessible for use on consumer devices.
- Quantized versions of the models are available, allowing for efficient local deployment with minimal resource requirements.
Questions Answered
What is the purpose of this presentation?
The presentation aims to provide insights into DeepMind's new models, demonstrate their usage in projects, and encourage audience interaction through questions.
What is DeepMind's mission and what recent developments have occurred?
DeepMind's mission is to build AI responsibly for humanity's benefit, with recent releases including new models and features like a computer use API and managed agents.
What is Gemma 4 and its capabilities?
Gemma 4 is an open model family with various sizes, allowing users to download, fine-tune, and expand its capabilities for diverse applications.
How can users interact with Gemma 4 locally?
Users can run Gemma 4 locally in their browsers, allowing for tasks like creating tables and analyzing data without sending information externally.
What are the performance characteristics of Gemma 4?
Gemma 4 models are designed to perform well on commodity hardware, offering cost savings and accessibility for users with lower-end devices.