transcribe

Leveraging Compute | Intel Corporation

AI Council · 25m · transcribed May 2026
More from AI Council Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 you [Music] I used to work in venture capital and unfortunately my most well known startup as a fake one so AI is particularly high fee as you probably know so during nips the big AI conference we did a we over breakfast one day me and my team dreamt up a company called rocket AI the technology was called temporally recurrent optimal learning spell out

0:31 troll but unfortunately people missed that and we made a website and we threw a party and we convinced all these journalists I won't name which publications that this was a real company we had seven very large investment funds for an invest in us and it took a little while it was when Business Insider kind of tried to write this this hit piece on us that we weren't a real company that I revealed it wasn't a company first but it would

0:56 everybody was just a party theme that uh kind of went off in one direction but again it's now my most famous association in the AI field and it was a drunk joke but uh I guess that says something about my career actually actually the best bit about it was not only what did the police come and raid the party which says something for an AI party but we saw t-shirt software's that said rocket AI with troll on the back

1:20 and I made more money off the t-shirts and we spent on the parties so yeah VC but I'm not here to talk to you about fake AI companies although there is some reference to AI hype in my presentation I I actually hate public speaking I think only sociopaths and academics like public speaking and that's probably a very fine line between them anyway so what I decided to do to try and get back into public speaking now that I am a

1:44 technology professional again is that um I wanted someone asked me to speak in an event I'll try and give it a title that's somewhat big enough that I can when I start making the slides a few days before which is what I normally do as I hope everyone else does and is that I have a big enough title that I can then think about all the things I'm interested in and turn that into a presentation so when where's reached out

2:05 to me I get the title leveraging compute which is definitely theme of this talk but I added last night the etc to highlight the fact that I have probably got ADHD and there is like 20 different things in this in this comic in this presentation so I'm sorry that it's not super hip may perhaps super Karen I tried Intel eight months ago I've never worked in a team bigger than seven and we have 10,000 people in our group 112 thousand

2:32 people around the world it's big it's intense like a country I started in architectural graphics and software cuz my background was working in AI investment and I moved to Silicon engineering because I just found it fascinating so I'm very very new to thinking about that and in terms of strategies like in engineering I'm not thinking about how the strategy in terms of what's on the silicon I'm thinking about how we should be designing computers and computer chips for the for

2:55 the future five to ten years out so that's kind of like the theme of this talk across my many varied interests so point number one I think computing has an optimism problem and here's an example why and I'm not just shilling for Moore's law because Gordon Moore is the founder of Intel I've always been interested in in Moore's law because Moore's law really is a self-fulfilling prophecy so Moses Gordon Moore states in 1965 that the rate of transistor density will double

3:27 approximately every two years and it was a really a prediction I think that's the wrong way of thinking about it it was more that he set a benchmark and then people thought they could go out and build it so to me this is like one of the most extremely powerful examples of what self-fulfilling and optimistic prophecies can do in technology but that doesn't mean to say that people still think is going to continue so this is I

3:52 wish I had it when more iPhones have a laser I guess that's very dangerous but um so here's some examples of kind of like tech cynicism especially around Moore's law I think my favorite one is in the top left everything that can be invented has been invented in 1899 lots of stuff about Moore's law dying across across time really like it goes all the way back to like 2009 there is nothing new to be discovered in physics 1900 people

4:19 are just really bad predicting kind of like the limits of progress which is a kind of another theme of my talk but they've been particularly bad doing that on transistor density and Moore's law so in terms of us safe social contract it's probably one of the most under Tomatoes social contracts this is the kind of multiples that we've had in Silicon engineering and computer chips and just between 1980 and 2020 so our group made this slide and you can see

4:46 that Gordon Moore was off by like a factor of a million so we really tend to push past what even people predict and I think the optimistic message about technology will be the thing that drives the future if we want it to be so what has kept Moore's aura alive for 50-plus years well you have more complexity because computers are getting more complicated you have abstraction at every layer layer you have innovation but the thing that I actually change

5:13 from the original slice this was a slider my boss Jim Keller who is amazing and you should fall in him he was probably the best chip designers the last 50 years and he he had a he had refactoring there but I changed it to belief I think people believing that Moore's law will continue is the thing that allowed more to continue because they believed that was the next benchmark for what they should design but most lawyers in just one technology

5:34 isn't just simple you know exponential growth of transit or not that simple that isn't just transistor gate transistor transistor density Moore's law is the kind of like name for what it is many many different innovations in in hardware that come together and not just in hardware but in physics and in math too that compile to allow for something that is the increased numbers of transistors and this is kind of what I get to learn about every day so when I

6:00 moved to work in the silicon engineering group the reason why I picked it was that I just got this amazing kind of like rapid PhD in all these areas of physics and electrical engineering that I would never have access to you I'm still a very nascent in my understanding of some of these areas but I wanted to highlight one and that is lithography which I'm kind of obsessed with right now so you know lithography is the how

6:23 you use lights to define process of features and by I think you can see in this image by 2005 they kind of like hit a wall in in progress at a hundred and ninety-three nanometers in terms of how much they could with the resolution of what they could do and it took 15 years to get to the current in they're using in quantity like seven nanometer designs now which is 13.5 nanometers and this is done by a

6:54 technology called extreme ultraviolet lithography and if you ever want to be amazed by what humans can do just in one second to watch a video and what happens in Eevee is unreal I fell in love with science when I went to go tour an Eevee fab but this is a data science /ai event why should you care I think the unsung heroes of AI progress have been datasets and environments and access to compute and if you look at a

7:24 lot of the kind of breakthroughs in AI it's really been these two circles that have allowed for those breakthroughs to be demonstrated and the person who really pointed this out in recent time was rich Sutton here a general read this paper especially a blog post called the bit of lesson by Richard Sutton raise your hand if you read it ok everyone read it it's so good and so rich sudden daddy of one of the daddies of

7:47 reinforcement learning he wrote a blog post in earlier this year that said you know if you if you look at progress in AI a lot of people spending time tweaking models trying to understand you know search and optimization but if you look at how the breakthroughs come it tends to be that something comes along and there's a new way to leverage compute and it was really powerful music for me to read this because I've even published one AI

8:10 paper but it was in reinforcement learning and then to read rich sauce and he's always been the hero of mine talk about how even in his own field the the progress has really come from computation justified why I had moved into hardware so so rich Sutton reinforcement learning here's a brief history of RL I'm gonna use this as a case study to kind of demonstrate the how much compute changes what we can do in AI so you know it kind of like stems

8:37 this kind of reward reward feedback stems in animal psychology and then it starts getting interesting again in one it's interesting for a bit I mean dynamic programming definitely paid to it but Richard Sutton and Dawson wrote the book about reinforcement learning and then the thing that really kept you know capital catapult catapulted RL was in 2014 when when deepmind used it to demonstrate their RL agents in in the atari games but again the unsung hero of

9:06 that demonstration to me was the arcade learning environment so we talked a lot about reinforcement learning in terms of the deep mind breakthrough but not that many people give enough credit I think to 2013 when the Atari games environments was created to allow for the agents to be demonstrated in that so there's one which is datasets and environments the second thing is on the other circle is compute so openly I made this great graph me cuz it meant that we

9:34 didn't have to make it at Intel so if you if you go back to Alex net the combinational neural network and you go all the way up to alphago zero you see a 300,000 X multiple in in demand for compute so that really goes to show just how much more compute that we're using for different machine learning paradigms so here's alphago alphago versus Lisa doll the last time I was in New York was three years ago and I was wearing a

9:59 t-shirt that was in sporting Lisa dolls so it goes to show that I'm on the human side so you know they had and I and I took this I was researching for this together date but um I think that there might have been some TP usage when they were using this alphago against Lisa doll but I wasn't quite clear but at least compute to the same level of around 1200 CPUs and 176 GPUs versus one man and one coffee computers is quite

10:27 cool so I really think need to start giving compute more recognition than then people have been doing and what's also interesting is the more I've kind of gone into hardware is look at how there's like this kind of symbiotic relationship between hardware and software so I'm thinking about AI working in venture capitals starting fake AI startups and then I go into hardware and I end up in this meeting which I just found fascinating which was you have this problem when you're kind

10:53 of doing a silicon design which is the amount of time it takes like design and test and validate the chip is is probably the lead time that you have to bring out the product against your competitors so you're always trying to reduce that time reduce a debugging time reduce the violation time and one of the way is that groups have been doing this this is a group of Intel the I've been working with is they use

11:11 things like reinforcement learning to actually debug invalidate so it's I found that interesting to think that there's like this relationship where computers allowing for the progress in the in the AI but then AI is being used to again reduce the time to debug back in the hardware ok next tangent section of my talk progress is weird why am I saying this here's Bell's law I think this is really interesting you kind of have these kind of arcs of compute classes and that

11:43 approximately every 10 years so you know starts a mainframe now it just because of Moore's law they keep getting smaller and more compact and and the reason I bring this up is that there's this kind of been running trope for the last few years that you know progress technology progress is really slowing down you know why is why aren't more exciting things happen and I want to make a counter example and in this presentation I've been thinking about this for a while

12:07 there then maybe we were focusing too much on the micro and not thinking enough about the macros because when you're in an architect knowledge yard it kind of seems like a diminishing return because you don't know what the next law of accelerating returns are going to come from so whilst you're in one of these little micro and curves on these micro s curves you can't see the overall kind of exponential growth of the s curves when you're in Moore's law if you

12:31 think about Moore's Law run transistors essence density you might have seen that lithography was going to reach this wall in 2005 and you wouldn't know that the arc of UV was coming so from from the inside from from the micro it can it can seem like it's ending but I actually think that progress can be counterintuitive on the macro and it and there's a reason to be optimistic so here's some of the s-curves of computing different programming languages over

12:56 time process design around chips you know now we have evie architectures each of these were curves that at one point someone thought they'd reached the end and something new comes along so I don't think that we should just be thinking that progress is slowing down and computing I think we should be thinking about what the next arcs are that are going to contribute towards the macro and here's opening I doing another graph me recently that again saved me from

13:21 doing some work they started plotting the usage of compute and if you if you see that they've tracked it's kind of like two year doubling that kind of follows Moore's law up to about 2013 and then it kind of goes up to a three point four month doubling they say and they kind of find this in these like two areas of computer and sorry two ears of compute and I think it's actually a really interesting slide because you think well

13:44 what happens in 2013 2014 that allowed for this kind of you know increase and I really just think it's money and they say this themselves in the paper which is that you know you start getting more money into a field there's more of a drive an economic drives I build better hardware to support the new startups and the new companies being brought out and I think that deepmind prepaid a large part in this because do you mind got

14:09 acquired by Google it started to show people that kind of research AI groups had a had a high value and it meant that startup money kept going I mean I prove this with Roc AI right we created a fake company and all these venture funds start invest in us where the money goes the growth goes to so you really got think about from the economic standpoint and Marie Shanahan from from England he made this comment about that the year

14:30 2014 which was a it was a year when a childhood passion became a global obsession and I really think that he looked back to the recent hype cycle of AI it does really go back to what was happening around there with deepmind's valuation and so economics pushes growth this is Jevons paradox this is actually applied to a fuel usage which is the idea that you know as fuel becomes more economical then we'll just use less of

14:55 it and actually what happens is that people just end up using more I think it's a same for compute compute will just get better they'll get cheaper you'll get faster but the demand for it will always outpace the the innovation in that sector next tangent does complexity kill progress I bring this up because I watched this presentation I'm an advisor to a distributed personal saw the company called other if you don't know check it out and born of the

15:27 engineers that sent me this talk this is Jonathan Blow video game designer the guy behind braid and it's a talk I think it's called a Banting civilization collapse has anyone seen it it's really good so I watched it like seven times it's really it really hit me and I came to the office and I said to everyone I've seen this talk it's the best talk I've ever seen and it kind of talks about how software is becoming more complex and we

15:51 depend on it and he's trying to show you that from the insight from the arc of the software from what she's observing in the in the s-curve it that it looks like that it's a it's it's on decline and he's got a point I have my own version of how I've interpreted it his talk better this is the lines of code in millions between different programs so you know version 1 of UNIX not that many total Google web services now as a 2

16:16 billion but that covers so much but there is a lot of complexity now in in software and again for the amount of executed lines of code per written code so then the kind of narrative becomes well software is in decline and our standards for software and getting lower and I think that's true but I think there's another way of interpreted to interpreting it which is that yeah complex systems are complicated maybe we need to get over it

16:47 - maybe this is the wrong crowd but we have too many expert Python programmers and too few generalists but that's okay because generality is tough and I try and learn as much as I can about how computers work at every stack and Intel and it's just endless computers are so complicated now that you have people who actually understand full systems I mean people that Intel they only know there

17:17 are a little one bit I mean some people obviously know more you'll get told off for saying that but but the idea that people could understand the interest in Trinity is of like the entire compute stack is is unreal now that as computers have gotten more complicated so then the argument becomes do we need a revolution in computing and you kind of see a doesn't there's another presentation that a Jonathan Blow referenced here and his his YouTube talk which is more in

17:41 his presentation which is the thing is called a 30 line code problem or what it's called I was trying to link to it at some point but uh it kind of presents as argument there's a lot of people saying like okay software's got complicated we need to go back to first principles and it makes me laugh a bit when I hear this some science because it kinda reminds me of Communist China and I spelt computer on

18:03 on purpose because I do this sometimes at work which is that this like should be like one computer it's gonna run more like you know we've just got this like one better instruction set that replace x86 and we could all just move together on this one paradigm and go forward and all be in agreement in a centralized or decentralized way I'd say that Erb it was going for that kind of decentralized way then you know progress would happen

18:26 a lot better but the counter-argument to that which which I'm gonna go into is that it's kind of amazing how much stuff we've done with all this legacy code and it's just like unreal to me that what we can be cynical about it that like you know it's bad it's complicated and we don't know how computers work it's also kind of amazing that we get by and we do it every day and it's also maybe kind of

18:48 up to us code quality is somewhat of the leadership problem we don't want to get our hands dirty and debug across the stack we're so used to programming there being problems and it's just wasting our time and you know maybe wasting our time is a bit of efforts like we used to software not working which is one of the comments that Jonathan Blow had it's a bit like democracy we don't just go and fix it we'll just keep compiling the

19:11 problems while compounding problems but there are other industries and there other complex systems that they don't have this as an issue for instance aerospace aerospace is a very complex system you have different vendors different providers you have many different layers of technology and lots of legacy code and in 2017 my mind that there was no eight highest years of a commercial passenger jet and that's despite there being more flights than than ever before and I think this is a

19:38 good analogy sometimes for programming and computer science as a whole that you know aerospace engineers just can't let it be so shitty because otherwise people die and it cost billions of dollars which is not the thing that we have to deal with every day when we're building remember building codes so my argument against the that software is is in decline is that maybe software is in decline but maybe the answer is that we've got to go

20:04 in and fix it and get our hands dirty and this was like the first mean that came up when I searched to get your hand study that wasn't pornographic okay so maybe we don't need a revolution in computers but maybe we need a revolution in a IBS raqqa day I was a good example of how ridiculous it was I think you were listed I won't name the company as one of the top 20 breakthroughs in AI in

20:31 2017 with a not real technology and people have been talking about AI yes for a long time here's actually my favorite AI paper it's written by Hubert Dreyfus in 1965 and he goes through he talks about different benchmarking issues and the problems with kind of having this kind of like narrative a narrow way of doing performant performative benchmarking and in machine learning and you know things haven't reach ange that much since 1965 so modern AI as we have both compute

21:02 access to compute and datasets and environments allowing us to create breakthrough demonstrate breakthroughs we're also able to hide behind them compute and datasets and environments allow us to demonstrate stuff that actually isn't that general tool that can seem general and here's an example and I mean I really like opening I but I just wanted to give the the dota 5 example as one is that you know notify train on 45,000 years of game play out of more than 100 characters in dota they

21:29 only had they only had 16 and then once the once it was like publicly available for non for like non champion humans to play against it very quickly humans were able to figure out how to beat it that's because bring in the general and these software agents aren't and it's problem is tech presses is quite stupid and they write these like flashy headlines I'm glad that I don't know if the New York Times CTO guy is still here New York

21:52 Times is ok and I wanted to give that credit Cade Metz is actually quite cool so but most the type press is just terrible and I know that because again I started a fake AI company and got pressed for it here's Barney pal he is an investor and a good friend here PhD thesis in 1993 which I love we was talking about how games as benchmarks for machine learning a pretty poor because the only thing that we do is

22:21 show that the that the program is able to demonstrate within a very narrow framework of a game it doesn't mean that it can generalize but however they'll present people will present it as if that there's some sort of level of generality so if you're interested in AI does anyone read this paper and it came out the other day no oh yes wise but I read this paper actually on the someone sentence me I read on the flight here

22:47 and I always thought that so France watch today is the founder of chaos but I was found him kind of annoying on Twitter so I never wanted to read it even though he only just came out but I read it and it just respawns everything I loved about computers and machine learning it's such a good paper I already hope that people read it he just it could be ten papers in one paper it's so darn brilliant he does a kind of

23:11 critique of past and present AI benchmarking so different things that people have used games and the different things that big tech companies have used to demonstrate their agents and their programs the other thing he does is he kind of goes into psychology he he redefines what he thinks as the is the right explanation of what intelligence should be in terms of generality and then for like math fans like like me he then uses algorithmic and face

23:33 information theory to quantify how to generalize how to generalize how to quantify the difficulty of the generalization like how do you how do you test for kind of like the broad skill sets and then and then at the end if that wasn't enough he provides a dataset a link to I get hardware you can kind of like test out your own agents in his new scale for for intelligence it really is a fantastic paper and he

23:58 reminds me kind of of the old AI paper so just don't really exist anymore so I'm jontron like there's 70 different things what are my main takeaways I think we need optimism in computing again and not just hype I don't think Moore's always said it's only dead if you want to believe instead and then we won't start building the technology that we need to make there be better pootis data and compute fuel AI progress but we also hide behind them progress

24:29 can look like it's slowing and declining at the micro but from the macro it can be an exponential growth we just can't see it as engineers people need to get their hands dirty and fix things and also create incentives for other people to fix those things should read shows paper and very much his point about algorithmic information theory is that a who here knows who Claude Shannon is love Claude Shannon so I don't think we have enough

24:57 generalists now but I think Claude Shannon is an example of an amazing journalist and actually he's like the father and his mate in theory but uh but like just like nai there are too few journalists and his Hubert Dreyfus in the paper alchemy and AI which I referred to earlier quoting something that Shannon said in 1965 and I love this quote so much because he said this not as a necessary hardware guy as you know already thinking of it as a

25:22 mathematician but he says can we design a computer and whose natural operation is in terms of patents concepts and vague so our similarities rather than sequential operations on tangent numbers and that's why I like to think about every day and I hope that more people who think about software and data science and and machine learning will talk more with people who are working in hardware because we really need better conversations between the two thank you [Applause]

25:51 [Music]

Summary

The speaker shares their unconventional journey from venture capital to the tech industry, highlighting the importance of optimism in computing and the interplay between hardware and AI advancements. They argue that while computing may seem to be slowing down at a micro level, macro trends indicate significant growth driven by innovation, belief in progress, and the symbiotic relationship between hardware and software.

- The speaker's background includes creating a fake AI startup that gained unexpected media attention.
- They emphasize the need for optimism in computing, citing Moore's Law as a self-fulfilling prophecy that has driven technological advancements.
- The relationship between compute power and AI breakthroughs is crucial, with historical examples illustrating how access to compute has enabled significant progress.
- Complexity in software development is increasing, leading to concerns about declining software quality, but the speaker argues for the necessity of addressing these issues rather than seeking a complete overhaul.
- The importance of datasets and environments in AI development is highlighted, with a call for recognition of their role in facilitating breakthroughs.
- The speaker critiques the media's portrayal of AI advancements, emphasizing the need for a nuanced understanding of generalization in AI.
- They advocate for better collaboration between hardware and software professionals to foster innovation and address current challenges in technology.
© transcribe · For agents Built with care and craft by Gokul Rajaram