Transcript
0:05 What I want to do today is give an overview of this class. I'm I'm actually curious before we get started. this is now September 2025. how many of you you know just started at Stanford? Raise your hand. Wow. Cool. Awesome. So others pay attention to who just raise their hands and do say hi to them and you know help welcome all the people they just joined Stanford in whatever program. CS230 is a class that we offer in the flipped classroom format. And what that means is that instead of listening to me or Kan my co-instructor that you meets that you meet next week, instead of listening to us lecture at you for an hour, R 20 minutes or whatever, we actually ask you to watch a lot of the video lectures online so that we can then make use of the precious in-classroom time for much richer, deeper discussions. So both today and for the entire quarter I would really warmly welcome anyone raising questions raise your hand and in fact I I I find that instead of you sitting there oh even though we have a longer session scheduled by the registra we'll use usually only up to an hour and 20 minutes for this course.
1:25 and the goal is to it turns out that a lot of Stanford students were watching you know the lectures on sego or on the online videos anyway. So rather than us delivering that lecture, same lecture year after year, we put a lot more effort to put very high quality lecture videos online and we'll ask you to just watch that online which people are doing anyway but just highly editive offline watching spend and spend the classroom time you know doing the things that make sense for us to get together in person too. so because you haven't watched any of the lectures yet or I assume most of you have not today we'll even maybe have a slightly shorter session to introduce the class talk over logistics and so on right as many of you know deep learning is one of the latest hottest trends technologies of computer science and AI if we look at you know say our PhD emissions or or even mast's emissions a very large fraction of all students that come to Stanford or applying to come to Stanford want to work on AI, right? I'm not sure I'm allowed to say the numbers, but they are, you know, extremely high, as you can imagine. And so, my goal and my my co-instructor Ken's goals, our collective goal is through this quarter to help you get to near or at pretty much state-of-the-art with regard to deep learning. and make sure that all of you walk away from this class highly skilled at applying deep learning.
3:00 and so it turns out that a lot of progress in AI over the last I don't know decade maybe 10 15 years was made by scaling and one of the reasons why deep learning was so successful was because it's good at absorbing a lot of data. So you know if I draw a figure where on the x-axis I plot the amount of data we have for a problem then using more traditional machine learning algorithms so you know logistic regression maybe decision trees using all the generations of AI machine learning algorithms as you gave it more and more data the performance or the accuracy of the more traditional algorithms would plateau, right? It was as if. So take speech recognition with older generations of algorithms, even as you fed it more and more data, hundreds and thousands and tens of thousands of hours of speech data, the accuracy would often plateau and it was as if the older generations of algorithms didn't know what to do with all the data that we now have. But what we start to find about 10 15 years ago was that if you train a small neuronet network also known as a small deep learning model this performance would kind of maybe get better and better. And if you train a medium-sized one and if you train a very large neon network the performance just keeps getting better and better. And I think the reason that deep learning has dominated the AI scene for the last 10 15 years is because there is a recipe for training very large neuronet networks that we can then shove a lot of data into that results in exceptional performance. So I think we started to see this because of some frankly some Stanford research papers about 15 years ago when when you know did the first early work on using cuder programming and GPUs to scale up deep learning. oh by the way actually has one fun fact the first my first GPU machine used to train neuronet networks using crew there which is a controversial thing at the time. It was built by a Stanford undergrad in his dorm room. Right. His name was Ian Goodfellow. But I think that compute server built in a Stanford undergrad dorm room. allowed us at Sanford to lay the early foundations of using CUDAR a time a modern language for training GPUs to train large neuronet networks and then you know obviously that influenced a lot of people and helped scaling up deep learning take off. So I I I tell that story because sometimes the work that some of you can do you know in a dorm room or graduate house student housing or whatever or in a lab at Stanford looking back over some number of years it can really have a huge impact and maybe in this class some of you will do work as impactful as that someday as well. But so what we start to find was that as you train larger and larger neuronet networks they could soak up lots of data and drive exceptional performance. and then there was a research paper out of bu which showed that as you scale up neuronet networks as you scale up deep learning algorithms the performance gains are actually quite predictable. So you can forecast if you buy this many GPUs throw this much comput and this much data at it what would the performance be? and then later open AAI popularized the idea with a really influential paper on scaling laws.
6:43 and that predictability of how deep learning gets better in performance then drove a lot of the investments in you know data centers and and building very large AI models with lots of data. and so let's see in terms of where this class sits. So in computer science all of us build on each other's work right a lot of the way that computer science and AI has made progress is we build on top of other ideas that are in turn built on top of other ideas that in turn they in turn built on top of other ideas. So maybe I want to give you a little map to to to show you maybe where deep learning sits. so I think there are you know machine learning is built on top of computer science. So I think it's actually helpful to learn CS fundamentals. And even though I use I suspect vast majority of use AI assisted coding be it tools like cloud code or Gemini CLI or open codeex or cursor or wind surf or whatever. I find that people that know CS fundamentals that really understand computer science that really understand how computers work rather than like I'm going to vibe code this you know like when you understand CS fundamentals you get things to work much better but on top of CS fundamentals there's a set of machine learning skills.
8:10 So how do you build algorithms that can learn from data? and then deep learning is a special type of machine learning. Well, that's really the most effective type of machine learning as far as I can tell in which we train neuronet networks. We train certain types of algorithms to learn from large amounts of data. so far in the last 10 minutes, you probably heard me use the words deep learning and neuronet networks. And I think today those two terms are almost interchangeable.
8:40 some some purists will insist on some technical differences. Raw practice doesn't mean the same thing. But what happened was the term neonet networks had been around for decades. But around 10 15 years ago a number of us realized that you know deep learning it was just a much better brand. and so even though neuronet networks have been around for decades, starting about 1015 years ago, it was deep learning that took off because who doesn't want learning that is really deep, right?
9:12 It's just a good brand. but you hear me use those terms kind of interchangeably, but deep learning algorithms, neuronet networks, they give us a way to take advantage of more and more and more compute capacity. So they can build very large AI models with a lot of parameters to soak up the large amounts of data to get more and more you know intelligence or to make better and better predictions or make generate more and more accurate outputs using the large amount of data that's available to us. Ever since I was a teenager my mom's been trying to convince me to stop mumbling.
9:48 But now many years later I still struggle with that. So I I'll I'll try my best. So please wave at me or or kind of let me know if I start to drift lower again. Yeah. All right. My I think my mother would be very happy that to practice like this now. all right. So CS fundamentals, machine learning, deep learning, and then the recent generative AI revolution. You know, generative AI sorry bad handwriting. generative AI which is mostly built by a specific type of neuronet network called a transformer neuronet network which you learn about in this class. to actually learn what is the transformer architecture later in this quarter as well is in turn built on top of deep learning right so I assume all of you you know are regularly prompting OM large language models to help you get work done and what I find is that while I use OM all the time for a lot of applications just prompting LMS it doesn't cut it there are a lot of things that I cannot get to work just by prompting an OM and so often have to go a layer deeper into the deep learning layer of abstraction and fiddle with deep learning algorithms in order to get certain things to work. And in fact, so what this class covers is we'll try to make sure that you are, you know, expert near expert in deep learning by the end of this class, but we'll also kind of dip a little bit into machine learning concepts. So we'll talk a lot about objective functions and tips and tricks for optimizing parameters in efficient way. and then we'll also actually reach up a little bit to cover some Gen AI. In particular, we'll talk about what is the transformer network. And then through this quarter, Ken and I will also chat a little bit about the job landscape as relates to Gen AI and deep learning and how to use how how deep learning is enabling certain types of Genai applications. Okay.
11:56 I have more to say about this, but let me just pause for a second and see if you have any questions. >> Yeah, go for it. >> Oh, good question. What I say prerequ congratulations. You've won the prize for the first question asked in CS230 2035. So, well done. Yeah. so is machine learning a prerect to this course? not really. So I think two common entry points to AI at Stanford are CS well a few common entry points are CS129 CS229 and CS230.
12:35 if you don't know any machine learning this course may end up going a little bit fast may seem like it's going a little bit fast in the first I don't know two or three weeks but some people do pick things up quickly that way. and maybe a screen. So a few causes that you may hear about. so 129 is a relatively easy entry point into machine learning that slow tends to you takes a longer time to go through the core concepts of machine learning like what's optimization objective? How do you implement gradient descent? What is logistic regression? What's a very basic neural network? So this is a relatively applied easiest of the of of of this list. CS229 which I'm also involved in co-eing is much more mathematical and theoretical very high pace very intense and very mathematical and this is less applied than 129 and 230 but this will go over a lot more of the theory and the math derivations behind machine learning algorithms. So for example, if you ever want to learn how to do calculus not using row numbers but calculus using matrices and vectors, you know, that's a bunch of sort of I don't know mildly complex math that's worth knowing. CS229 goes over that. CS230, this class is relatively applied and it focuses just on deep learning. So of all the machine learning, there are a lot of machine learning algorithms out there, right?
14:08 learning, unsuized learning, a lot of machine learning algorithms out there and many many of them are very useful but and and so they're all worth learning about but of all the machine learning algorithms out there the one category that is most useful is you know that's taken off the most is deep learning and this class focuses just on that but the other arms also worth learning about. So 129 is the easiest on ramp. but if you've done either 229 or 230, I would probably skip 129 at that point. But if you are getting started there there are multiple on ramps.
14:39 >> Yeah. >> Can we take 229 and 230 together or do you think it's about >> Oh yeah. Can you take 229 and 230 together? Yep. Thank you. Price for the second question. Feel want to clap. Sure. Go for it. yes, you can take CSU30 and CS29 together. we designed the two curriculara to be relatively low in overlap. So the very small amount of of overlap between these two, >> right? All right. We we won't clap for every single question, but yeah, go ahead.
15:13 >> are you going to cover more like recent learning algorithms that are used in recent lab? >> Yeah. So will we cover the recent OM developments? We will touch on the transformer neural network but not the latest OM variations in this course. I think a lot of the it turns out that when you go out to get a job well maybe right if assuming you go and work in industry the number of people training LMS is actually very small right so those jobs seem to be incredibly well paid. We hear about kind of very high salaries in news media, but the vast majority of application builders end up sometimes working at the geni level, not that often trading a transformer from scratch. but then often using deep learning tools as well.
16:09 So maybe one example something that many of my teams have done is we have trained transform model foundation models from scratch but relatively small ones in in say startups. But one thing we do do quite a lot is take a pre-trained transformer network and then engineer our own data to further fine-tune it. Right? sorry if I'm using words that you may not totally understand. You know pre-trained fine tetun you know what all those terms are by the end of this quarter. So those are things that we actually do kind of day-to-day this is important for getting a bunch of brothers work. So you will gain the foundations needed to do type of work in this course. the one thing one thing we don't do is talk a lot about how to train the largest cutting edge transformer networks. I think that is a very important skill set is a relatively niche one for which some people are you know getting paid really really well but the number of people doing that in the world is actually small whereas the number of people building applications with you know this set of skills is is is very large. Do >> you know any courses that you cover?
17:16 >> do I know any courses covering that? I think Percy Leang was thinking about doing something but I don't remember he's doing it this quarter. few people are think thinking about doing something like that. Yeah. >> this is a question about the course itself. do you know what portion of the course is going to involve like mathematical analysis and what is coding? So so I'm just repeating for the for the mic for the home viewers. What portion is coding versus what proportion is mathematical analysis?
17:48 This course is relatively math light. sorry maybe that was too strong a statement but I think this this course is very practical. I I I I remember many years back I was speaking with a mathematician. and you know we were just chatting and he was asking he was just talking about his career why he chose to be a mathematician and I I still remember you know this he had he had like stars in his eyes when he told me that he chose his career path because he felt his role is to pursue truth and beauty in the universe and that's why he became a mathematician.
18:28 In this course, I'm not going to do any truth and beauty stuff, right? So, truth and beauty is good, but you find that I I want to take a very practical approach to, talking about how to build applications and build software that works. yeah, cool. Anything else? Right. Cool. Awesome. All right. Thank you for all the questions. Please keep them coming and feel free to interrupt me or Kan throughout this quarter as well with questions. I love it. so, oh just to just to flesh this out a little bit more, this is what I see in terms of teams building practical applications. And I'm excited about applications because with improving machine learning algorithms, deep learning algorithms, Jenny algorithms, there are a lot of applications that you can build now that just were, you know, impossible or or like really inaccessible, you know, to to any person to build even a few years ago. and so I find that when I prompt Gen AI, it works really well for a lot of textbased applications. and there's work on multimodal LMS, large multimodal models.
19:44 So making inroads into vision, making inrows into audio, but really genai algorithms, especially transformer networks trained to output text, you know, like chipd claw, gemini, and so on, really fantastic for textbased applications, right? and I find myself regularly working with deep learning algorithms directly when I am working with audio data, image and video data. and then also a lot of structured data.
20:22 sorry my handwriting is awful. Right. So structured data refers to large tables of numbers right like giant you know Excel or Google sheet spreadsheets but so that's that's structured data unstructured data refers to text audio images maybe video and because a lot of genai right language models like chat GBD had grown up being text in text out kinds of machines they are remarkable for a lot of text processing applications but for other types of data I end up often and you know dipping down directly to use various deep learning algorithms. and then it turns out that for textbased data if all you do is prompting you could usually go quite far. So a lot of applications are built by prompting OM but but I've been on you know quite a few teams where after fiddling with the prompts for a month you just can't get the performance to be better just by tuning the prompts.
21:25 or another good problem to have. It turns out use of geni tools are relatively inexpensive when you're prototyping, right? It's like you know a few dollars per million tokens. You can do a lot. but sometimes if you're lucky enough to be on a product that his product market fit and a lot of users want to use, you know, multiple times I've been on teams where we basically did not care about our launch language model bill, right? It was like whatever, you know, $20 a month or $100 a month was fine. But when more and more users start using it, then to your team's kind of positive surprise, your AI bill really starts to skyrocket.
22:05 And then at some point, you look at how much you're paying for your AI bill for the large language models. And to bend the cost curve back down, often lot of the techniques in deep learning become very relevant as well. So there are I don't know, I'm thinking just some of our bills were really breathtaking. I don't want to say the numbers, which is kind of definitely more than we want to pay. You know, as much as we love the companies providing OM, our bills that were paying them were significantly larger than I wanted to pay. and then knowing how to use deep learning to fine-tune smaller models. That was really the critical skill set that just bent the cost curve back and just made the whole thing affordable to keep on providing a service. Right. So, yeah. Right.
22:50 So that's what so that's the skill sets I hope you get from this course. Let's see. All right. that was right and to give a sorry use this one.
23:21 So to give a f to to give a quick overview this course the online materials is broken down into five modules. So just give you a overview of the five of them. first one is on the basics of neuronet networks sorry NN neuronet networks and DL deep learning right. So you learn how to build a neuronet network or deep learning algorithm from scratch in Python. I find that sometimes if you use the frameworks like TensorFlow or PyTorch it hides a lot of the details.
23:50 So we actually work through how to build a basic neuronet network and how to build a basic deep learning algorithm just in raw Python. So you really understand it. and then the second module, the second miniourse will be on how to improve, how to tune your neuronet networks. So you may have heard that when you're training a neuronet network, there are a lot of parameters or we call them hyperparameters. Hyperparameters are parameters that control the parameters, right? So the weights are parameters. Hyperparameters are things like the learning rate or what's the size of your neuronet network. And so there are actually a lot of hyperparameters that we end up tuning and try to give you a way sense of what are the most important ones and practical skills for tuning them. It turns out that if I look at you know like my PhD students I think every one of them that right well I think I think that probably every one of them definitely everyone that every PhD student I know that became great I think at some point wound up at 2 a.m.
24:55 tuning hyperparameters right and I still have very clear recollections of like being in the office you know 2 a.m. 3:00 a.m. filling with parameters, try to get it to work. And it turns out that literally your skill at tuning hyperparameters, it really makes a difference. So there were some evenings that I knew my skill at tuning hyperparameters, you know, frankly, it made the difference between whether I went home to sleep at 3:00 a.m. versus where I went home to sleep at 7:00 a.m.
25:22 Maybe this is not maybe don't do what I do. I'm not encouraging this type of behavior. but but this is just my personal experiences but it really makes a big difference your practical skill at how quickly you can figure out the recipe to get these neuronet networks to train. and and related to that is one thing we've chatted a lot about in this course is strategies for building machine learning projects.
26:00 So it turns out that if you build a complex system you know let's say you build this is one example we'll go through later this quarter. Say you want to build a system that recognizes your face, a camera that recognizes your face, your friend's face, unlock a door, right? Security, safety, whatever. Build. So system like that, you know, I I've worked on systems like that. Chance work that these are complex systems with multiple components. there's a camera there, you know, do you subtract clean up the image? Do you cut out the face?
26:29 How do you register the face? How do you compare a face? How do you decide to take another picture before you unlock the door or just you know or someone trying to fake hold up a picture printed on a piece of paper? There actually a lot of decisions. And so what I find is that the biggest difference between a team that knows how to drive forward a project like this well and get it done in days rather than weeks or weeks rather than many months is the ability to drive a disciplined development process. It turns out when you have a complex system, less experienced teams will often almost pick things at random to work on, right? They'll read one research paper and say, "Oh, I read in the research paper we should get more data." You know, well, because like some newspaper said AI needs lots of data.
27:16 So, let's go spend six months to collect more data. Turns out a lot of the time collecting more data does not help your application. but sometimes it's a huge help. So given your application, how do you decide? Should you spend more time collecting data? Maybe you should buy more GPUs. I I actually definitely know people that read in the news a lot of GPUs are helpful, right? And then I've literally met, you know, fairly senior business leaders that have spent a very large amount of money buying GPUs and then, you know, some funny stories and then and then I go talk to them and say, "What are you doing with these GPUs?" And then sometimes, you know, there was one meeting I was in where literally a a very large familyrun business had bought a lot of GPUs and the CTO then pointed to his nephew who who who was a who's a current college student undergrad and said, "Oh, my nephew knows AI. I'm giving him this very large budget in GPUs and I think he'll do AI for me." Right? And so, but so I think that knowing how to make these decisions and not just buying into the hype that you read about in the newspapers is is really important.
28:32 and one thing I hope to do in this course is share with you what driving a disciplined development process looks like because this is one of the things that really makes a 10x difference in the speed with which you can get something to work. I I I've literally seen teams, you know, like take six months or 10 months pursuing an approach that experienced engineers would go in and go, you know what, I could have told you six months ago that spending all this time collecting data, this was not going to get your application to where you wanted to go, right? but so how do you examine an application and figure out the diagnostics to figure out what are the productive things to do for your application? We actually spend a lot of time on that. And in fact, I'm excited about doing some simulation exercises in this classroom with you later this quarter where, I'll invite you later this quarter, you know, to to to say in this in this scenario, what would you do and and see if you could, you know, make decisions in a more systematic way, right?
29:33 all right. then course four, we'll talk about con convolutional networks. very useful for computer vision applications. and then so connets are specialized models most you used mostly for vision applications dealing with images and then we'll talk finally about sequence models. So sequences could be time series or sequences of text like words. to also touch on the transformer network that you know that power a lot of the gen revolution. Okay. and so throughout learning through learning all of these things I hope that you gain a large tool chest that enable you to tackle an almost bewildering range of applications. I think one of the things I've most enjoyed as a AI person is it turns out when you work on AI, there are so many other teams that have data and that could use our hope, right? So I feel like as an AI person, you know, I somehow bizarrely had a right to play in, you know, building autonomous helicopters or hoping companies, I don't know, place more relevant ads. maybe not the most inspiring thing I've worked on but certainly very lucrative for some companies right or improve web search rankings or improve safety you know get rid of the kind of negative toxic results you may not want search results to come back on or improve e-commerce retailing or improve speech recognition or help ships be more fuel efficient right there all these are real examples or fight fraud which is actually really exciting when you're fighting financial fraud which is obviously a bad thing but when you sometimes you when you're fighting fraud, sometimes you wake up in the morning and your team's alerted you that there's a new scam and then you just have to go and fight them and build algorithms in real time and you know that every hour you take you know more more money actually least of the financial system. So it's kind of awful that this financial fraud but it's actually one of the most exhilarating things I've worked on because literally every hour you are slower to to you know formulate response you know more dollars are leaking out every hour. So anyway so so somehow when you have this tool set of deep learning you just have a bewildering right to play or ability to play in a huge range of applications that could use your hope right and I think on campus too there's so many departments across campus in the sciences engineering humanities business that have interesting data where your skill set will let you if you choose collaborate with them to do interesting projects and and I find for example bizarrely some of my PhD students are working on climate science it's like what do I know about climate science right I wish I knew more but using machine learning tools you know we could actually work on climate modeling and and geoengineering and and just play in all of these important important hopefully important and interesting places I hope that you have that skill set as well by the end of this quarter. Yeah.
32:52 >> If you have enough data for network versus like it's better suited for other like machine learning. >> Yeah. So how do you know if enough how do you know if you have enough data for neuronet network? it turns out to be really difficult to know. so if it's a application that others have worked on or that you've worked on before then you may have a sense. So for example, I know I because I worked on face recognition, you kind of have a sense if you want to train a face recognition system from scratch, having like 50,000 images, 50,000 unique faces is pretty good, right? Or or so if you worked on it or if you read the research literature for something people have worked on, that would give you a gut sense for how much data could be enough to get you started. but then for green few brand new projects that no one in the world has worked on before if you can't find parallel projects are kind of comparable it could be really hard to tell and so common advice for completely green field sorry green field I mean a brand new project dissimilar than things what so for example if someone has invented a new medical device and no one has collected this type of you know blood specimen data before it's really hard to tell and in that case the most Common advice is get a little bit of data and just try training a model and the degree to which your initial model works or does not work that will help you hone your perspective on how much data may be needed and you may be surprised maybe 100 data points is all you need right sometimes we've been surprised by that and then sometimes we've also worked on applications where you know 100 billion data points later we're still trying to get a lot more data and I I find it really difficult to Good question though. Anything else?
34:43 Oh, so this is a good time for me to pause and take questions because so what what I'm going to do is right what I'm going to do after this is switch tracks and talk a little bit about exciting trends in AI kind of recent trends in AI that I'm excited about and how I view the AI landscape. But so this is actually a good break point to see if people have other questions before I talk about some trends in AI. Anything else?
35:13 >> Yeah, please. >> When you were differentiating between those different tasks, what were like the disting? So I guess let's see all of these many of these terms are blurry and kind of fuzz a little bit into each other but when I refer to generative AI generative AI is this body of work that generates text and sometimes also images sometimes also audio using deep learning algorithms in certain ways. So I think geni refers to this body of work with most of the center gravity on generating text you know maybe also images and the text generation algorithms have been mostly implemented using transformer neural networks train on large amounts of data you know straight off the internet and elsewhere so so when I refer to genai that's I guess that's one particular application of deep learning models that has given us large language models like traiki and cloud and gemini and meta the llama and so on. Does that make sense? Yeah. Cool.
36:29 All right. Anything else? Cool. All right. So, let me could we go to the slides, please? One of the nice things about this class is Ken and I can occasionally just share a few things we're seeing in the broader world. Hey, while while doing that, I'm actually curious how many of you use a specialized AI assisted coding tool like cloud code, cursor, geni, codeex, curts.
37:01 Awesome. Almost everyone, but not everyone. Interesting. Huh. Interesting. Oh, I thought interesting. Okay, cool. So, yeah, one of the most exciting things that's happened in programming is AI assisted coding. And I feel like I personally hope I never ever have to go back to coding by hand, right? Like it's just it's actually interesting. I often work in coffee shops on the weekend and a few weekends ago I was sitting in a coffee shop and sitting next to me was someone that was you know coding by hand and looks so strange. I just asked them what are you doing in a in a in a in a respectful way you know and it turns out that the that they're they're doing a homework from some other university that required the code by hand but one of the things I find exciting is that individual programmer productivity is much higher than ever used to be right and maybe I want to share with you just just one one one thing what what I see so I find that in the software work that I do I I maybe categorize it into two buckets. One is building quick and dirty prototypes to see if something works and then sometimes I you know write production grade enterprisegrade robust reliable software that has a scale right. and I find that where AI assisted coding has made the biggest difference is building the quick and dirty prototypes.
38:27 whereas I think actually literally one of my collaborators used one of the agent decoders that I named but I won't say which one but literally this morning he sent me a slack message saying sorry you know this agent coder had a database migration error and we just wiped out all of the database records for fortunately for a test application with like five users but but it did happen. So I find that oh thank you. So I find that my use of the agent code is you know for the production grade software is more careful whereas for building quick and dirty prototypes it kind of so you're not shipping software in respon in irresponsible way.
39:08 quick and dirty prototypes have lot fewer dependencies. you usually don't need to integrate with legacy data infrastructure. And then, you know, I'm I'm I'm going to say something that I that that feels like something I'm not supposed to say, but I but I'll say it anyway, which is I find that when I am running code, I I find that people often worry about safety and reliability of software or security of software. So one thing I often say to my teams is if you're building a prototype that only runs on your own laptop and doesn't use any sensitive information so there's no risk of sensitive information then unless you are planning to you know maliciously hack into your own laptop right the security requirements can be lower and so I find that when building quick and dirty prototypes having a sandbox environment that lets you operate within it quickly means means that you can just make a lot of decisions faster without worrying as much about scalability or security or reliability. So long as a sandbox environment means this stuff isn't going to get out there or leak information or create a security loophole. So that that that's part of what lets us move much faster.
40:24 And so I find that because of the speed of prototyping so you can do so in a responsible way to pursue innovative ideas. my teams were increasingly, you know, try 20 things and see what sticks and because I and and I know that a lot of teams are lamenting that many proof of concepts never make it into production, right? You try something and it doesn't work. And I know some teams are feeling angst about that. I actually have a different view. I think if the cost of a proof of concept is low enough, then who cares if you have to do 20 of them and that's a price for finding the one or two things that works really well. So one thing you hear about in this course is both when when you're building a machine learning application, you usually don't know what's going to happen. and there's specific reason for that. The reason is the output of a machine learning algorithm, it depends both on the code you write as well as on the data you're training on. And while you control the code 100%, you don't really know usually what's really in the data in the weird and wonderful data that the world has given you. So for example work on speech recognition a lot in multiple companies in multiple contexts and even now when I work on speech recognition I'm still you know sometimes a little bit surprised that oh this data has people of a certain accent more than I realize or people somehow speak faster or boy there's a lot of background noise when people use it in a car right so even though I worked on speech recognition multiple times oh and actually one recent example application I I was actually surprised by the number of background speakers. They talk to us, they come and talk to a different person, then they talk to us and then we get confused, right? So, so I find that the data that the world gives us is often weird and wonderful. And so it is only by building a system that you then start to discover these things in the data that lets you make progress. and with a lot of software applications as well separate from machine learning applications a lot of what I end up having to discover is what do users actually want right so again I control my code 100% I can write whatever code I want I control that but you don't get to control how your users will react to your system and I find that our ability to build quick and dirty prototypes rapidly both to discover what's in the data and also to take the users to see if they like it. That allows us to drive faster feedback loops than was ever possible before to then help us build more and more valuable software, right? and I think you know I I I I know that the mantra move fast and break things, right? Got a bad rep because it broke things. and I find that some teams took away from this that we should not move fast, but I think that's a mistake. So, what I usually tell my teams is move fast and be responsible. And despite all the hype about you know AI extinction risk with AI kills all that somewhat bizarre hype in my opinion I find that when teams move really fast we can then implement things test out in a responsible way and much more quickly identify problems and fix them. So, I find that a lot of the most responsible teams I know, teams able to really get stuff to work really quickly, they tend to be some of the fastest moving teams because it's that speed of execution that lets you finally implement it, figure out what's in your data, figure out what users want, and that's the best way to figure out what could actually go wrong to then make sure things don't actually go wrong, right? and somewhat related to that is AI coding assistance. and I I I assume you know almost everyone or everyone in this class knows how to code. if you if you are well you haven't learned to code yet, you probably might not want to take this class yet. but I find that there have been people including very senior right business leaders advising others to not learn to code on the grounds of AI will automate it and I I want to share this with you not because I think you need to learn to code but because I hope you help me spread the word right go to all of your friends in other departments to tell them this advice to not learn to code I think we look back on this as some of the worst career advice ever given and and the reason is when coding becomes easier more people should do it rather than fewer.
45:05 So when humanity went from punch cards to keyboard and terminal that made coding easier and so more people learn to code. when we went from assembly language to modern well at that time modern programming languages that made coding easier more people learn to code. I I went back I actually found these papers on when these articles on when Coobo very old school programing language right was invented and there were actually people that said oh wow now we have the coobo programming language coding is so easy who needs programmers anymore right and and obviously the opposite happened from text editors to then you know AI assisted coding as coding becomes easier people should code a lot more a lot more people should learn to code. And the other thing I'm seeing is I I just add just something that's been on people's minds. I know that unemployment of recent computer science graduates has ticked up. It's higher than it's been you know compared to I think the last decade. and so I know that to people learning CS that has caused some conststonation. So I'm going to share my view on that. So it turns out that what I see in Silicon Valley and beyond Silicon Valley is we just can't find enough people with these skills, right? I know many businesses that would well I know large businesses that love to hire a thousand people with skills in GI and all deep learning and machine learning but that are struggling to find people with these skills.
46:40 conversely, there are still, you know, universities with curricula that has not changed since has not changed much, right, for the last like, I don't know, since 2002 before the rise of Genai and so I do see that many new CS grads, not from Stanford, but from, you know, around the country are struggling with finding jobs because unfortunately that all their nonAI enabled skill skill set that is you know not as much in demand right and maybe just for myself heck today I will not hire someone I I will not hire a software engine that doesn't know how to use AI to help them with coping just it just doesn't make sense same reason why I just won't hire someone that uses a punch card instead of keyboard and terminal right and and I think when the world evolved from punch card to keyboard and terminal people still had punch card jobs for a while but eventually the punch card jobs just went away it just doesn't make sense anymore so today There are still coding by hand jobs around maybe some some specialties some very low-level coding where AI is not very good. Turns out AI is not very good at sometimes with GPU programming.
47:48 You know, there are some niches with coding by hand actually still makes sense. But for building applications you know so I I I remember a few months ago now many months ago now where I interviewed two engineers back to back one had not yet graduated from college but was highly on top of geni coding. So you know spoke with that candidate knew how to use AI built proc quickly. Right after that I also interviewed someone with 10 years of experience as a full stack engineer but whose skill set was exactly the same as their 2002 skill set had not tried out any AI assisted coding really good skill full sack engineer with 10 years of experience. And it was actually really clear to me I picked the fresh college grad where well he had not he was about to graduate over someone with 10 years of experience. Right. So I think making sure you master these skills are really important and what I'm seeing is there is a very large gap that businesses are having a hard time filling for people with these skills.
48:58 but the demand for the 2022 skill set software engineering full set engineing skill set that is not there right so so I think and then in terms of AI assisted coding I find that CS fundamentals really are important so in addition so I know I I I applied someone fresh college grad over someone 10 years experience that true There's actually one other part to this story which is with with with respect to all of you about to graduate from college. The best programs I know are also not fresh college grad. No disrespect intended to fresh college grads. Some are about to graduate from Stanford. the best programs I know are really on top of AI assisted coding and additionally deeply understand computer science fundamentals. Right? So it turns out maybe I'll illustrate this with a quick story. when I was teaching an online course my team wanted to generate background pictures like this you know just for decoration. So when I was working on this this is of course generative air for everyone. I was working with a collaborator Tommy Nelson that understood art history and so my collaborator knew the language of art. He knew the artistic genre inspiration, the pallet, so he could prompt midjourney AI image generation with the language of art and so he could generate beautiful pictures like these.
50:25 In contrast, I don't know art history. I wish I did like and so all I could do was go to mid journey or AI image generation and I type, please make pretty pictures of robots for me. And I could never get the control that my collaborator Tommy could to generate pictures like these, which is why we use all of his pictures and none of mine. Right? And I'm seeing the same thing in computer science.
50:57 one of the most important skills for the future is to understand how computers work and understand how genai and deep learning and machine learning work so that you can use the language of AI, use the language of these tools to tell a computer exactly what you want so the computer can do it for you. And there is actually a huge difference in performance between you know someone that's learned to just prompt an LM without understanding how computers or how AI really works versus people they can look at the problem analyze it and then with AI assisted coding you know tell a computer how to take the next steps. which is why I think that CS fundamentals is very valuable. CS fundamentals, machine learning fundamentals, deep learning fundamentals. We my I and my teams we use that knowledge like every day, right? In in making pretty consequential decisions, right? So, so I hope that you get that from this class and the many other classes at Stanford as well, right? all right. I think I might leave the rest. There's more I could say on trends in AI, but I find that AI assistant. Oh, but but one thing I hope you do really, you know, go to all your friends in all the departments across campus and encourage them to be a builder because the other thing I'm seeing is clearly for computer science professionals, you know, use AIS coding, no CS fundamentals, build cool stuff, but for other disciplines as well that is not computer science or not AI, I'm finding that you know the education professional or the climate scientist or the mechanical engineer, the ones that know how to build software are just more productive and get a lot more done.
52:47 And the barrier to entry to AI to coding is the lowest it's ever been, you know, in our lives. And so this is a good time, frankly, for I I I wish every single Stanford student, right, will learn to build software with AI assistance. so I hope you go help your friends across campus to master those skills as well. Okay. Oh yeah. Any any other questions? Yes. Go ahead.
53:21 >> AI. I've heard a lot of like a lot of the industry folks talking about how they would rather hire someone with 10 years of experience who knows just AI as is like just using coding rather than fresh undergrad that does come like that deep understanding of deep learning machine learning and like AI as fundamental just because they have more experience. Do you see that also? >> Yeah. So let me rank productivity right and I'll give you four levels I think and again and I say this with a lot of respect for individuals so if I talk about productivity is not with any disrespect or any lack of affection for anyone or for their work right but I think you know least productive are people with no experience and don't know AI right one step on top of that is people with less experience but on on sorry one step on top of that is people with say a decade of experience but that don't know AI. on top of that, I would rather take a fresh college grad that does know AI, but then even more productive is someone with, you know, a decade of experience and also really on top of AI. So, I think between the two factors, really understanding AI is very important, but experience is also important. And so the best developers I know we just work and we just ship code like was no one's ever done I think even two three years ago are very experienced developers.
54:52 They're also very on top of how to use the latest AI technologies please. Oh oh sorry and just one other thing about the job market. I find a lot of employers have not yet figured out how to hire appropriately. This is contributed frankly a lot of employees you know if a company has no one that knows geni how do they even know how to interview appropriately so that is a problem that we need to solve as well.
55:15 Go >> ahead. Like for example I'm thinking of classes like CS 1071 where we're not encouraged to encourage. Is that true of the fundamentals? that's >> yeah boy. So CS17 CS11 are great. so so do take them if you're considering them. I find that the fundamentals are important.
55:49 how do I put it? Yeah, I'm not sure what else to say other than that. I think I think honestly Sanford were known for CS department were known for really I'm I'm I know I'm biased but I want to say like probably the best entry level CS program classes of any university in the world. I'm pretty biased so maybe shouldn't say that but I think they're excellent course if you want to learn the fundamentals in really solid way and I know the instructors are routinely thinking about how to update the curriculum and realities of geni. So I I think they do an excellent job with that mix.
56:23 piece. >> How do you define a person know comparable to person just out the person? >> Yeah. How do I define someone that really knows geni compared to someone that just use a tool and typic I feel like in Genaii there are two buckets of skill. again by the way deep learning is not geni deep learning is also a very valuable skill but syni I think there two things. one is I find it really useful to know how to use AI coding assistance that's really valuable but having you know kind of fundamental knowledge helps right you do that and then the other thing is in genai there's a number of emerging tools I'm going to toss some buzzwords okay so if you don't know what any of these words mean don't worry about it but but I think there are emerging tools like rag retrie generation or how to vector databases how to do evals and error analysis how build guard rails.
57:17 how to use knowledge drops and interface that with your OM. maybe how to do multimodal, how to fine-tune a model. what else? I'm probably blanking on something. Oh, how to build agentic workflows. But I feel like that these categories of new techniques built on top of Genai that are like a useful bag of tools for building applications. So when well frankly when I'm interviewing candidates you know which still do a fair amount of is is is for geni row is these skills I I tend to look for this set of tools as well as AI assisted coding and and then yeah please >> I don't know if this question is best poss to you if I'm considering taking 229 230 at the same time both of those classes are final projects so how would you recommend thinking about ideas to pursue yeah, so let me just give one tip about CS causes in general, not just 229, CS230, which is I encourage you to think of AI causes at Stanford a bit like Pokemon.
58:23 You got to catch them all. But in all seriousness, I think taking more CS causes in AI, it it is a good thing to do. Definitely encourage you to take multiple causes. I think in some years we've had students do joint projects where the standard is higher higher expectations for sure but that that that's that's one option to to look at. Yeah. >> All right. Go for it.
58:59 >> Oh so let's see. 229 and 230 are very different causes. 229 well in 229 you we have live instructors doing the lectures in person rather than online content so it's not the flip classroom so in CS230 we have most of materials prepared kind of in online videos highly edited online videos and then C I think the big the other biggest difference is CS229 is more mathematical theoretical but there's important math right whereas CS230 is more practical ical and so I don't know if we do even one a single maybe we do like one proof somewhere in the lectures but we just don't do a lot of math in CS230 and it's very a lot of machine learning is very empirical you should try see what works but have a disciplined approach for exploring what works we focus a lot more on that in in CS230 and 229 covers a lot more techniques right so there a lot of machine learning techniques siz learning learning you know I don't know decision trees boosting ky's clustering. So CS229 covers a much broader survey of a lot of machine learning techniques where CS230 is just one thing. Well, deep we go really deep into deeper