transcribe

Webinar: What AI Can and Cannot Do: Intelligence Augmentation in Practice with Michael Bernstein

Stanford Online · 59m · transcribed 1h ago
More from Stanford Online Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 Right. So, we would like to thank everyone for joining us for this live session today on what AI can and cannot do, intelligence augmentation in practice, brought to you by Stanford online in collaboration with global alumni. Please note that this session is being recorded and we're delighted to have Stanford faculty Michael Bernstein with us today.

0:30 Michael is a professor of computer science at Stanford University where he's the Bass University fellow and senior fellow at the Stanford Institute for Human Centered Artificial Intelligence. He's a nationally a nationally best-selling author and a creator of the G of generative AI simulations that represent the highest cited research in the history of UIS. Michael focuses on designing social, societal, and interactive technologies. This research has been reported in venues such as the New York Times, TED AI, and MIT Technology Review, and Michael himself has been recognized with the Alfred P. Sloan Fellowship and Computer History Museum's Tech for Humanity Press. Michael holds a bachelor's degree in symbolic systems from Stanford University, as well as a master's degree and a PhD in computer science from MIT.

1:22 Today, Michael will explore how organizations can use AI to augment human intelligence, evaluate AI opportunities, and create lasting value in product decisions. To ensure the best quio audio quality, please make sure to stay muted throughout the session. And we encourage you to turn on your cameras so we can make this an engaging and interactive session. If you have any questions, please make sure to share these in the Q&A box and we'll address them when we get to the Q&A segment at the end. If you are looking at the bottom of the screen and you do not see the Q&A option, please make sure and select the three dots and you will see the option there. So with that out of the way, I will hand it off to Michael.

2:01 Thank you so much for being here with us today and the floor is yours. >> All right, it is wonderful to see you all. I am excited to dig in and to talk a bit about what AI will can and can't do for us. when you're thinking through your AI strategy, what is it that you ought to pursue and this is select as as Lisa mentioned these are selections from you know an online course that that I offer and I think that it's been resonating with a number of you know company boards executives senior executives CTO's you know CEOs and so on that I've spoken with and I hope it resonates with you as well. So let's start from the question we're trying to answer. So what does it feel like as time is going on in terms of the amount of AI progress that's going now?

2:59 You're probably tracking that there feels like there's a new model every week. In fact, there was a new model last week. It might be gone today, back tomorrow. We'll find out. the but the broad question here is like what does it feel like? It does really does feel like we're on some sort of rapidly accelerating curve. And we could debate exactly the shape of this curve, but it certainly feels like it's moving very very quickly. So a second curve here, same period of time. What does it feel like in terms of the number of AI based products and services that are getting launched? And I if your feeling is anything like mine that I think it's probably you know at least as as as steep at this point that people are really trying to figure out what is this useful and useful for and like what can we apply it to help solve problems that our at our organizations or that our our customers or people in civil society might care about. So then a third question time is going on. All right.

4:02 How many successful AI based products and services are we seeing? And I think here the curve is a little bit more anemic. You might say, look, there are absolutely many many different ideas being thrown around. It's a Cambrian explosion of ideas in AI and you know, people are going to experiment. A lot of stuff's going to fail. That's kind of by design. But from what I've seen, there are legit differences that can separate out the successes from the failures. And I want to spend this time together kind of piecing apart the why here. What is it that might be separating these successes from these failures?

4:48 And in particular, what do they know that others don't? And why are so many products misfires, abandoned, embarrassing, you know, wastes of money? What are you supposed to do about it? And I want to split this my answer to this question into two parts. One is around how you decide what's feasible and the other is how you decide what's worthwhile, what what you should actually build.

5:22 So let's start with what's feasible because AI is moving pretty quickly arguably too quickly for our you know old school business strategies and this is because you know if any change you want to make can take months right you you make a plan you you have to or execute an organizational change you have to actually you know do the do the work you have to prep the launch actually do the launch you know manage manage the change with your customers and so on. And at the same time as you're doing this, the AI capabilities change. So by the time you actually launch or go to market with something, the technology itself has now shifted and what you launched looks old. So what do you do about this? What I think we need is not some sort of answer that says, "Oh, hey, I'm going to tell you exactly what's going to happen 6 months from now." I could do that, but I think that it would be out of date very quickly. So, what I think we need is actually like a durable framework that's not tied to specific technical road maps, but lets you plan ahead. And that's going to be the the the promise of the first part of this presentation.

6:40 And what I'm going to do first is I'm going to lean into a a distinction that I found to be quite useful for planning. And that is a distinction between what we're going to call sharpedged and roughedged problems in AI. So these are I'm going to make a claim here that is the framework and then I'm going to explain what the heck I mean and then we'll sort of walk it through to explain why I think it's so powerful. The claim, which you're not going to understand because I'm not going to have defined any of my terms yet, is that generative AI struggles more on something that we will call sharp edge problems than on something else that we're going to call roughed edge problems when you try to deploy it. And it's not because the AI is better at one or the other. It's be, you know, the model is equivalently powerful. It has a lot more to do with how we embed those and deploy them. So think the claim is going to be about harder to to launch something sharpedged than roughedged. Okay. So what do I mean? Rough-edged problems are any problem that has a ton of different solutions that we would all consider to be correct. I'll give some examples in a moment, but just to sort of give the rough intuition here, you know, if you're having an AI, you know, write something like write a marketing copy or, you know, write a summary of this lecture, or draw an image or play Mario. Anything where getting 80% of the way is still helpful. Anything that has many different possible solutions typically also has the property that if you get me sort of in the neighborhood, it's still an accelerant and I can get and I can take it the rest of the way.

8:29 In contrast, a sharpedged problem is a problem that has only one correct solution, one and only one. And what that means is that if I'm if I haven't hit exactly the right solution, then it's just dead wrong. So we'll talk about agentic tool use agents and so on in a moment. if you expect your for example your your agentic software development environment to you know be entirely bug free on the first shot that's basically a sharp-edged problem. Only one way to nail that. traditional AI decision or prediction problems like you know is this person going to get readmitted to a hospital in the next next month that you're either right or you're wrong.

9:17 Anything where getting 80% correct is still actually just wrong. when you if you've ever heard the phrase almost only counts with horseshoes and hand grenades, this is someone who's discussing sharpedged problems. horseshoes and hand grenades I guess they would argue are roughedged. Everything else sharpedged. Okay, so that's what rough edged and sharp edge problems are. And you can look around and you can see sort of where AI is working in each of these different spaces. So video generation is a rough-edged problem. You know, if I have some goal like make this video of a man riding a horse in the American West, there are a million different videos that I would all look at and I'd say, "Yeah, that seems pretty plausible. That seems pretty good." so there that one's roughed. There are lots and lots of possible correct answers.

10:09 Now, we're talking a lot about agents these days. Let me let me talk about what that quite means. cuz they tend to sort of compound both sharp edges and sometimes some rough edges as well. And an an agentic AI, just if if you're not if you haven't come across this yet, is something that will output tool commands. So you take your AI, you'll call it an agent, and you give it various tools that it can use like it can write code, it can read a file, it can run some code, it can search the web, it can send an email, it can support, you can submit an order, it can open an application. you Agentic tools have access to this library of things that they can do and then the AI decides at any given time point which of these do I do. I watch what happens and then I and then I keep going. So if I'm trying to build something, I'm gonna say, what am I gonna do? Okay, I'm gonna write some code. All right. Okay, now I'm gonna go read these files. Now I'm going to go write those codes and so on.

11:06 And you know when when we empower an AI to decide which of these commands it can it wants to run and when, we refer to it as an agent. Now agents are interesting because you can look at how they're deployed in both rough rough and sharpedged problem capacities. you know, a sharp-edged version of a coding agent would be, you know, hey, fix this bug, ship it to prod, meaning, you know, there's a problem, fix it, deploy it. I'm either right and I've like, you know, correctly patched the bug or I'm wrong and the bug is still there or it got worse or it takes down my server.

11:42 Rough-edged version of this would be something where you're having the coding agent maybe try to add some logs for issues, maybe guess or hypothesize what might be going on, open a poll request, you know, to to investigate. These are things where there are lots of different ways to get to get it right. it's not just sort of a you're right or you're wrong. And in fact, you know, what we can see with coding agents is that you can actually wrap many sharp problems into inside of a coding agent as long as it's automatically verifi verifiable if it got it right. So, as long as it knows that it's wrong, coding agents can in some cases tackle sharp edge problems.

12:19 but if not, like if if you're just having a coding agent go off and like try to, you know, improve quality of life of or design the interface or something like that, then it's it doesn't really have a signal to to iterate on. Speaking of design, design is a rough-edged problem. there are many many different plausible designs for a given spec, but we often sort of treat it as sharp. Like we say, oh, I want a design for this. Like make me a signup modal and add a way to store user accounts and it's got to be right. Like we sort of assume we can just ship off a description and and get something back.

12:53 It's a it's a little interesting because like it's not like you would go to an architect and say, "I want a house with three bathrooms. Go." like it's often you know requires back and forth. So that's why I would describe it as a rough edge problem. There are like many many plausible interfaces given a given given a specific in input you'd give it. But we often sort of treat it as this like do this thing and you're wrong if you get it wrong.

13:19 So when does AI fail? Why do I say that AI struggles more with sharpedged problems than roughedged ones? I just want to reinforce it's not that AI is inherently better or worse for rougher shock merge problems in the sense that it's like technically more performant for one versus the other. again, it is the same model underneath. It's actually much more about us and it's about the it's about our tolerance for error.

13:54 And the observation is that our tolerance for error is often very very thin, very low in a sharp edge problem. And so think of something that is a sharp edge problem so that you're either right or you're wrong. And then examine like how good does this need to be for you to be able to trust it and use it in your in your everyday business life. Think of it this way. Let's say AI gets you know on the left side is it's AI is not very performant. It's not very accurate. On the right side of this graph, AI gets more and more and more accurate. When we are looking at the y-axis, we are considering practical usage like would this actually get adopted?

14:36 And as the when we're talking about a rough-edged problem, as the AI gets better, it sort of can get me closer and closer to the thing that I actually wanted and my my adoption will will rise. So, when AI gets me like a decent draft of some copy that I need to write, I'll use it and I'll tweak it. If it gets me better drafts, I'll use it even more. If it gives me an even better draft, I'll use it more than that. But that's not the case with a sharp-edged problem. A sharp-edged problem, you're either right or you're wrong. And so, the way that it often gets treated in practice is this very sharp edge, quite literally, where below some threshold, it's too too errorprone. you can't use it. You can't throw this thing in front of a customer.

15:23 Like, let's say that it's trying to do customer routing between customer service functions. If it's just too errorprone, I'm sure you've experienced being on the phone with some with some automated butt and it just keeps doing the wrong thing. You get frustrated and you hang up. Below some threshold of accuracy in sharp edge systems, you just you dump it. And then above that threshold, you can actually build it into some inner loop that you can just trust and rely on. But it's going to depend on where that tolerance is. And honestly, for many sharp edge problems, I think those those tolerances are much lower. Like that is to say, we need we make much higher expectations for automated loops than many AI systems can provide, especially once things just start being being accurate. So let's take an example. Spam filters. Spam filters have been around since the 1990s. Naive Bay algorithm has been around for a long time. And you know, it is something where honestly it has taken something on the order of 25 years before we can really start trusting them. It's this is a classic sharp-edged problem. it's either spam or it's not. And you either put it in my spam folder when you didn't when you weren't supposed to or you didn't put it in my spam folder when I didn't want to see it. And if you rewind just a couple years, I don't know if you were like me, but many people had to keep checking their spam filter daily because maybe there was an email from their boss that got filed in spam and it's just like, oh god, I didn't I need to make sure that there's some important there's no important email that I missed. And now it's much less often that essentially that accurate, you know, our our threshold for accuracy hasn't changed, but the AI has gotten good enough that finally you can sort of rely on it more or less and check it check the the spam filter less off less often. Another old school example is Siri or Alexa.

17:21 This is another sharpedged problem where you say, you know, to your to your I I worry about saying this out loud, you know, Siri, you know, play the Beach Boys or something like this and it will it either does the right thing. It plays the Beach Boys on your, you know, music or it doesn't and it plays something else or it sets a reminder or whatever. It's either right or it's wrong. And because there are so many errors in this right now, we often just retreat. we're like, "Okay, I can't use this for anything interesting. I'm just going to use it to set timers." Again, in principle, once it's accurate enough, you could rely on it to ask all sorts of things, but it's a it's a it's a series of sharp edge problems. If if it's if if I if it doesn't execute my command exactly correctly, then it's more problematic than had I just tried to do the damn thing myself.

18:09 And the trick is when you're using agents, these sharp edges can really compound. Think of a a game of telephone where this agent is trying to to watch, you know, is trying to watch u your your the parts in your warehouse and then order more so that you're the agent is responsible for just never making sure you're never out of supply. And so there are a bunch of sharpedged problems that are embedded there. Find the find the parts at the warehouse. Decide whether to order more or not. Pick how many to order and send it to the right ship it to the right warehouse. If any of this of these steps misfires, the entire workflow fails. But bear in mind that they are all sharpedged, right? So, you find like finding the part status at the warehouse. What if you look up the wrong part? What if you look at it in the wrong warehouse, right? and get anything wrong and the whole thing is is is borked. If you decide whether to order more, maybe your threshold is wrong. Maybe you think that you need more of these when you actually don't.

19:12 You've wasted money. pick reorder volume. You get too many, too few, you're not going to you're not going to have the supply you need. If you sh set delivery location, if you send it to the wrong location, there's only one correct location. Any any other one, if you ship it to Florida for the wrong, you know, when it's supposed to be going to California, it it's it's actively harmful. And so, it's like a game of telephone. sharp edge, sharp edge, sharp edge, sharp edge, sharp edge. If you have to be able to trust all of them simultaneously and this is why we see, and I'll talk more about this in a moment, you know, coding agents that are working on sharp edge problems tend to succeed when there are automatically verifiable outcomes because that way they can know when one of the sharpedged links has failed.

19:57 We see sharpedge failures all over the place. Legal hallucinations. You've probably heard of this one where a lawyer gets gets sanctioned or fined after they submit court documents that include a site a legal citation to a case that doesn't exist. Right? This is sharpedged. You're either citing the right support for for your argument or you're not. And there's no there's no in between. This is either, you know, a like a a real citation that supports the point or it's not. And so we see here someone who didn't catch it getting really embarrassed, right? And again, when we see sharpedged with motion with AI agents, here's a paper from Arvin Narayan Narayan's group at Princeton where they're studying reliability of AI agents. And again, we're seeing that these agents are not always consistent and robust.

20:56 And it, you know, has yielded, as I'm just quoting here from the bottom of their abstract, recent capability gains have only yielded small improvements in reliability. And that's because if you can't get each of these these elements over the threshold, you're really still in a world where where the whole macro thing is below your sharp-edged threshold. Looking at customer support bots, you can think about like DecaGon and Sierra as dueling startups in this area. You have really a a sense of both sharp and roughedged opportunities there where there are some sharpedged problems you could have AIS work on with with customers. Where's my order? You know, common troubleshooting, routing people to the right page, arranging returns.

21:44 These are all things that you can sort of succeed with in a sharp-edged in this. I think that most of these things are good enough at this point to to rely on as sharpedged. And then there's a bunch of roughed stuff like suggesting responses for human support or answering customer questions or a sort of concierge style interaction. These are all things that are more roughed. They're like many different ways to get that right. And so, it's we see them deployed, successfully even if the AI is not always 100% exactly the best, it's still typically good enough.

22:19 You can sort of think of this as producing feedback in different ways. Like a rough-ged problem might be able to give you a lowfidelity version that is good enough for you in your work to carry forward. Whereas a sharp edge problem might be closer to something that's high fidelity where it's like very has to be just shippable as is. And you know one one final example here would be something like accessibility feedback. Here's a a paper from from Germany showing how for blind and low vision people AI can help you understand where your UI might be missing accessibility guidelines. Right? So if you're a creator of software, you run it through this AI and it will help you see what what you might do. And this kind of thing can be helpful with rough-edged problems, right?

23:16 But often I get asked, okay, rough versus sharp, how now help me predict the future. That's what you promised, right? Here's the thing to know. The set of sh of solved sharpedged problems with AI is always going to be smaller than the set of rough-edged problems that AI has solved. That is to say, whatever that ball is of problems that AI can solve that are sharpedged that it's successful, there is going to be a larger ball of rough-edged problems that AI can solve today.

23:55 Why is that the case? Well, any problem I can solve to 95% or 99% accuracy, I can also solve to 80%. Right? And the set of things that I can solve 80% of the way is always going to be much larger than the set of things that I can solve, you know, 99.999% accuracy, which is what you might need for for a sharp-edged problem. So, think of it this way. When you're trying to plan ahead, know how tasks evolve in AI.

24:26 As AI gets better, a task is going to start, it's that yellow circle there is going to start outside both of these circles. It doesn't work at all. Right? But as AI gets better, it'll move from unsolved to solved as a rough-edged problem to eventually solved as a as a sharpedged problem. It's going to sort of like migrate inwards in this graph. And the reason it can't just teleport to sharp edged is because if it if if you could solve it in a sharpedged way, you definitely have a good enough model to to solve it in a rough-edged way. So then think what's going to happen when the next generation of AI models comes out.

25:09 Well, when the models get better, what's going to happen? Well, the set of things that we can solve in a sharp-edged way will grow, right? that inner circle is going to expand because the AI model has gotten better and more things it can solve really really accurately. At the same time that roughed circle is going to expand too. So things that used to be outside the rough edge circle are now going to get absorbed by that. So you don't need to have some any technical expertise in this to understand what the next generation of AIS is likely to make possible for you. You can just look and say is this problem just outside that inner border of you know maybe it's just outside that sharpedged border. It's it's definitely rough edged now and it's like barely sometimes good enough to be sharp edged but not good enough to rely on. Well, then the next ma, you know, model version, it may get sucked inside of that sharp edge, and you'd be well positioned to be ready to to leverage that. Likewise, something that like just barely works or barely doesn't work in a rough-edged way right now, the next model release probably going to suck it in and it's going to be good enough for roughedged work very soon. And you can be ready to go.

26:32 But you also don't have to just wait. You can often turn sharpedge problems into roughed edge problems and work today. People often start when they're envisioning AI, they start by picturing something that's a that's a sharp edge problem. Oh, I want a this predictor. Yes or no, I want it to automatically do this or that. You know what? You can often turn that sharp edge problem into a rough edge problem which will work earlier. You don't have to wait for the AI to be autonomous. For example, maybe you have you're building some AI predictor of hospital readmission. Maybe that's making too many errors for you to just rely on. But instead of just having it do it completely autonomously, you can have it write a report, a report that sort of lays out the risk factors to a human decision maker. That now is a rough-edged version of that same problem, right? It's like, maybe it's not good enough to do it completely by itself, but it is good enough to lay out some stuff that that I ought to pay attention to that I can use to make a better decision.

27:35 So, what's your rough edge, sharp edge strategy playbook? Start by by reflecting on your expected AI use case. Is it a sharp-edged problem or is it a rough edge problem? If you're envisioning a sharp-edged use case, ask yourself, what's the threshold? What level of accuracy would the AI need to have in order to be useful? How many nines in that reliability estimate do you really need in order to to to, you know, actually use the spam filter, the Alexa, the the agent, whatever it would be.

28:08 If today's AI cannot solve the problem at that level of accuracy, ask yourself, can I turn that sharpedged problem into an equivalent roughed problem where it is good enough today? That's a trick. If you're instead envisioning a rough-edged use case, then turn around, go out, you know, spend 20 bucks and try to use today's models and see if it can produce something useful as a starting point for iteration. This allows you to say, " I can make use of this today and I know where things are going because I can see that trajectory from from rough to sharp over time."

28:52 So, I set out by asking this question like, "What separates the successes from the failures? Why are so many products abandoned?" And I've answered half of that so far. I've said this is why this is what is possible what is impossible but there's one last piece here that I want to talk about about why you what you should build why you should build them and that has to do with the re the reason being here is that like you can have the best AI in the world but if it doesn't match people well it's going to get abandoned and I can show you these examples from from various robotics labs it makes it super visceral right where you look at something like these robots, they're not the most efficient robots in the world. Like if if these robots wanted to point at things and grab stuff, there are much faster and more efficient ways they could do it. But instead, they're, you know, using gesture, they're looking at you, they're they're sort of swinging their arm. you can the the human plus AI combination is much more efficient than if it did the thing that just the robot would have wanted to do. So something here has to embed the human- centered element.

30:07 And we have to think about this carefully cuz you know in this driving simulator where there's a self-driving car in a simulator, the car correctly says emergency, tries to eject control to the person. They barely notice in time, almost drive off the road, veer into the opposite lane, almost get into a car crash, swear, and like, you know, almost crash twice in within, you know, 30 seconds within this simulator. And there's something here that's going on about how even great AIs can produce huge issues.

30:40 This is something an example of what I would call and what we call in the research literature the seam the this the handoff between the AI and the person because of course you know a more advanced AI something that could have handled this as a sharp edge would have effectively just navigated these traffic conditions fine. And likewise, had the person been in driving control the whole time, they also would have navigated fine. The error wasn't at the AI qual AI, and it wasn't just with the person.

31:12 It was at that handoff, the seam between the person and the AI. And I have this this hunch that this is a widespread issue not just with self-driving cars but with many AI based products where my colleague Eton Adar likes to say don't let your user interface write a check that your AI can't cash right like that we're overpromising we're promising sort of a sharp-edged experience when all we can execute is like a roughedged a roughedged outcome and it's it's almost like many companies are putting engines on the outside they're saying like hey buy my AI powered device and it's like what about it is AI powered what problem is it solving and it's like no no don't pay attention to that it's AI it's like it well what exactly and I think this can lead to what people talk about as the you know the trough of disillusionment where you kind of overpromise and then everyone gets inevitably disappointed in what is actually possible because they're more because the the the push is more tech first than problem driven.

32:21 I think a much better way to think of this is something that's that we refer to as intelligence augmentation. And if you want to remember this term, the best way is just to remember that intelligence augmentation IIA is just AI backward, right? intelligence augmentation is a reaction to this idea that AI is just going to replace us. And it is a recognition that that's often both short-sighted and wrong.

32:53 Where a traditional AI would be to say replace human intelligence with artificial intelligence, what in intelligence augmentation or IIA would say is how do we augment our intelligence with AI? How do we give ourselves superpowers? How do we make ourselves smarter, more strategic, more clever, more funny, whatever, whatever it would be. How do we, you know, have a strap-on cortex that makes us do things that we couldn't do before? not trying to remove us but make us superpowered.

33:27 And in fact over half of usage of large language models today is already happening in an augment an augmentative capacity. The media loves to focus on replacement as the narrative and to be sure there are there is replacement that's happening and some jobs will shift. However, the the vast majority of AI usage right now is or maybe not the vast majority, excuse me, the majority of AI usage right now is in in augmentative capacity where people are using it to iterate on things, rough edge problems like you can't pull me out of the loop.

34:06 And you know, this is not a new idea. This actually goes all the way back to this Turing award winner Doug Engelbart who came up with this notion of augmenting our intellect all the way. You know, he in he used this to invent things like the computer mouse, the bitmap display, interactive word processing, collaborative word processing. All in the 1960s, this guy was a beast. Like it was really impressive. We augment instead of replace because often replacing does not go so well.

34:35 here's some research from my colleague Angel Kristen at Stanford finding that as these kinds of data tools get deployed by you know web journalists, legal professionals and so on, it is a it is it gets pushed back on if you just if people think that you're there to replace them. They'll drag their feet. They'll game the AI. They'll openly critique it be like oh you know it made this mistake that cost us a million dollars. you really should not like trust this AI.

35:07 And in fact, we have these issues where in order to make these things actually get embedded successfully in organizations, you often have to make it less less AI focused. here's a research project out of Carnegie Melon finding that to get doctors to actually use an AI, what they had to do was make it seem quote unremarkable. In this case, doctors famously protective of their autonomy and expertise. Well earned. They they have MDs. I don't. I'm a different kind of doctor. they will essentially they had to take the AI and kind of move it into the corner. Like rather than the AI being front and center for for the doctors, having it sort of be a little vis visible like floating thing on the side was saying sort of implicitly, you know what you're talking about. Just by the way, in case you're curious, here's some other information, right? It was not threatening the it was not saying it was not taking over. It was more invisible or less visible, I suppose.

36:11 But it's even more complicated than this. The goal, right, should be human plus AI does better than human. That's the entire vision of intelligence augmentation. Like if human plus AI didn't do better than human, then why are we even here, right? And we do see that this can happen. We call this in the research literature complimentarity. It means that human plus AI does better at this task than human alone or AI alone. And we see a number of instances where this does happen. My colleague Eric Bolson at Stanford has run experiments finding that for example at call centers you actually do see augmentative effects where people who get advice from a you know call center employees who get advice from AIs on how to handle really tough customers you know do do better. Likewise, here's a tweet by Eric Toppal pointing out how a large recent AI randomized control trial in medicine found that you know they have massive improvements in detection of breast cancer from mimographies and this is human in the loop. The AI was helping decide what to what to prioritize and so on.

37:18 BCG the consulting group found that large numbers of people of their own consultants improve performance when using AI for free creative tasks. This is all complimentarity. This is all also unfortunately not the complete story. if you look at the rest of this BCG article, it does say that on a different task where people took the AI's feedback at face value, their performance was 23% worse than not using the AI at all. 23% worse. And then they're not alone. You know, here's an RCT and software engineers finding that engineers think they're going faster with coding tools at the time, but are actually going slower. another report suggesting that many generative AI pilots at companies are failing.

38:08 What's going on here? Well, some researchers at MIT started to to investigate this. They did what's called a metaanalysis where they took a bunch of studies and then turned them all into one mega study where they were looking at this the patterns across them. And they drew this graph. Zero in this graph means that in that in a in a given study, there was no effect. Human plus AI did exactly as well as human alone or AI alone.

38:33 To the right in the green zone is a positive impact. That's complimentarity zone where AI plus human did better than AI alone or human alone. To the red to the left is negative complimentarity so to speak a lack of complimentarity where human plus AI did worse than human alone or AI alone. And the interesting thing about this if you look at this graph is that there is a lot of red. In fact, on average, people plus AI did worse than people alone or AI alone.

39:04 And I think that this is really interesting. What's going on here? It turns out that when they looked at what was in the red and what was in the green, the biggest losses tended to be decision-making tasks like will this person get readmitted to the hospital? And the biggest wins were sort of content creation generation tasks. To me, this is really interesting because this very much splits along sharpedged, roughedged lines. Many of these decision-making tasks are sharpedged problems and many of the the content creation tasks are roughed. So, we are definitely starting to see a like a difference in the actual on the ground success of AI along rough and short-edged axes.

39:51 So if our goal is human plus AI does better than human, but the reality is that that's not always guaranteed. We call that not great. Why? What's going on? And this will be my last point before we turn to some discussion. It's actually about human psychology. One thing we know that happens in AI is what's called over reliance. Over reliance is sort of this feeling of like we I'm flying where someone starts to use an AI and then doesn't doublech checkck it enough. So at some point the AI inevitably makes a mistake especially on a sharpedged problem and you don't catch it like the lawyer who didn't catch that they submitted a hallucinated citation.

40:35 So, this happens because people just go into a mode of satisficing or cognitive offloading or what some folks at Wharton call cognitive surrender where they just totally like kind of check out and trust the AI to just do the thing. And it turns out that even if the AI is explaining its reasoning, we don't always catch it because it turns out that paradoxically when an AI explains its reasoning, have you ever asked chat GPT a question and it writes you a whole essay in response? you're like, "Oh, that seems plausible." The more it explains, the more plausible it sounds.

41:08 So, after you overrely, you see that it makes a like a a really bad mistake. And then you pivot over to what's called algorithm aversion. Algorithm aversion is where you underrust the AI because it turns out a team of as teams of psychologists have found that our trust in AIs is more brittle than our trust in people. given the exact same error. Like if you get into a car with me and we drive to dinner tonight and I get into a minor accident, nothing, you know, you're not hurt, we're all fine. Would you get back into me into a car with me, you know, maybe not tomorrow, but okay, soon enough, maybe.

41:49 If you got into a self-driving car and that self-driving car got into the exact same minor accident, would you buy the same brand of self-driving car? Maybe not. That's algorithm aversion. Our trust in algorithms is brittle. And so people start by overrelying. They see that it makes a mistake and then they swing over to algorithm aversion and they underrust the AI. So if I pull this together, modern AI developments are opening many new opportunities for product development. I like to think of these by classifying them into this space of roughedged problems and sharp-edged problems to help us understand what opportunities are likely to open up when.

42:38 But even the smartest AI based product ultimately has to solve a real problem for real people or it's going to get dropped on the floor. And that's ultimately the answer to the question I laid out at the beginning that you can use intelligence augmentation as your litmus test to answer the question of what separates the successes from the failures. And I think if you do so, you will be much more likely to take strategic moves in your plans around AI development that will are going to pay off. So with that, I would like to pause and hand it over to the global alumni team who have a few words to say and then we're going to have some discussion Q&A until the end of the session. So let's hand it off over to global alumni. Thank you.

43:33 Thank you very much, Michael, for sharing all of your expertise with with this whole group that was able to make it today. if you have any questions, as I mentioned at the beginning of the session, please do use the Q&A box and enter your questions there. while we give you a few minutes to enter your questions, we wanted to just briefly mention that you know some of the topics that we covered today in this session are also covered in a course that we are glad to offer with Michael Bernstein. this is a course on AI powered product innovation. It's a sixweek course fully online with two faculty sessions as part of the of the course learning journey.

44:18 we'll give you a quick glimpse of what the key takeaways are for the course so you can get an idea of what that is. so key takeaways are identifying where AI is appropriate to apply and distinguish viable use cases from high-risk applications. we talked a little bit about that today but you can get more in depth on that with the course content. You can also evaluate how AI can augment human capabilities and support better decision- making. understand the psychological principles behind trust, adoption and human AI interaction and how they influence AI enable product decisions. address ethical and societal risks which which is very very important early through structured approaches to decision- making. Use generative AI agents and modern AI systems to stimulate user behavior and guide product strategy and identify what drives AI product adoption and long-term user behavior and a hands-on capstone project. so you can see here that there is a project that allows you to put everything into practice and think really through how you might apply all of the concepts that you learn through the course. we'll go to the next slide so you can quickly take a look at what the what you'll learn through the course. So when you walk away from the course, there are four main things that we hope you walk away with. The first is being able to apply structured frameworks to guide AI product decisions. The second one is to assess where AI can deliver reliable outcomes and where it cannot. And finally, anticipate how AI capabilities may evolve in the coming years and how this impacts product strategy. And then finally, make more responsible decisions as an outcome by recognizing ethical and societal risks.

46:05 And then finally, a quick walkthrough of what is included in the course. so beyond the content, you will also learn with AI tutor, which is a tool that allows you to ask questions throughout your learning journey, whether it's on the content or logistical matters. It will accompany you throughout the learning journey to be able to ask questions in real time. you'll be able to apply AI products frameworks through hands-on real world assignments and the capstone project as I mentioned. you'll engage in guided discussion forums and with your peers across the industries which will allow you to network and get to know the other people in the course. You will work with modern language models to explore intelligence augmentation and generative AI agents and you will also earn a Stanford online certificate of achievement upon successful completion.

46:55 so with that, I will hand it back over to Michael who will address the questions that are coming into the, chat box right now. and if you are interested in the course, we will make sure and drop the link so that you can go and get more information. Michael, back to you. And I see that there are some questions coming in. is there one that you would like to address just at the top of this section of the of the webinar? let's see. I see a question from Vanessa about whether you need to be an expert to critically analyze the results of the content especially around sharp edge problems. That's an interesting one. So I think that the question is what is the level of accuracy that you need? So for for example let's take a coding agent.

47:52 There are some things that these coding agents can solve in a sort of sharp-edged way that you don't actually need to be an expert at anymore because it's so accurate that you're probably like you don't your expertise may not be be may not be needed. So for example, if I were going to be writing a relatively straightforward algorithm to I don't know pass some variables around or count things or etc etc the models are good enough that despite the fact that that's a sharpedged problem it's either like the algorithm is right or it's wrong you probably don't need to be an expert to oversee it because you can just sort of trust it in the same way that there are compilers in my computer that are very complicated in how they optimize things and my programs or there are you know in a programming language they can do things called garbage collection that make sure I don't like leak memory and crash computers that I don't need to track I can just sort of trust that they work and this is because they're over some sharp edge threshold if it's not over whatever that sharp edge threshold that you need is you absolutely have to be an expert in order to watch it the problem is that if it's not over that threshold and you have to watch it, you start running into this AI over reliance problems where sometimes it's just too easy to go along with it. It's sort of like if you buy a a used textbook and then it's already got highlights in it.

49:20 It's hard to unsee the highlights even if you know they're wrong. Like this person was a bozo. What are they highlighting? But you still you still see it and it still influences you. And so even with expertise, you can get kind of caught. that that's tough. The the only ways around that would be that if you can use your expertise to create automated checks for the conditions of success that can help or you use your expertise in a more rough-edged capacity like when I am using a coding agent I have a PhD in computer science. I know how to program allegedly and yet I am still engineering when I use these coding agents because they're not good enough to just be completely autonomous.

50:08 I need to be like why did you why did you architect it that way? Why not that way? It'll say oh good idea, you know, very sickopantic and so on. And so I think you do need your expertise in the loop in order to to to succeed at these kinds of things. let's grab another one. I'm seeing some in the chat and some in the QA. So, I'll do my best to kind of jump back back between them. let's see.

50:37 Oscar asks around the metrics that you see as being important when evaluating success. I think that often the metrics that people think of first are are metrics that are derived from something being a kind of replacement solution. this thing that I do you know 10 times a you know 10 times a day I now need to only do once and you know that's great. the issue there is that we're kind of that only allows you to measure things that are transformed by replacement. And if you follow my argument, that is honestly, you know, an optimization that will, you know, save 10% or whatever. It's not going to create be the thing that creates whole new industries. And so you really have to measure what you actually care about. Like if you're really creating something intelligence augmenting, you have to track other metrics. Maybe instead of is it cheaper or saving time, it's more like are we shipping better products? Are we shipping are are our products doing better in the market? Are they crashing less often? Whatever it would be. And so I think that often what you know a like a what the seauite or the board wants to see are like clear ROI metrics and the easy metrics to make are typically replacement ones but they're often the wrong ones. So I am typically looking at okay if we're looking for something augmentative what or who are we augmenting and how would we know if that's working? That's the metric I push for.

52:30 >> Let's see. There was a question that came in about risk tolerance which might tie in well to what you were saying from Dara. so when it comes to risk tolerance is there any data on how successful AI solutions have been in customerf facing situations? So I guess could be a good followup to success metrics and how much tolerance do we have when it comes to different risks when something is customerf facing. Yeah, I think this is exactly a sort of rough sharp edge distinction, right?

53:01 Where our risk tolerance is often quite low when you're deploying an AI in a sharp edge situation. I mean, again, just think back a few years to when you call, you know, the airline and it's trying to talk to you and you're just saying like, I need to change the following flight and then eventually you just start yelling operator, operator, operator, right? like there's just the there's very we have very thin trust margins. This is what algorithm aversion tells us, right? And so as a result, you have to think through how often am I willing to let an error slip through on my system? And even worse than that, you have to say knowing that it's not like a random error that might happen occasionally.

53:51 It's going to be clustered because the AI is going to have consistent sort of classes of errors it makes all the time. Like, oh, whenever someone comes in with this particular kind of task, it's going to do worse. And so this is really a strategic decision at your business level where you have to say how much are we willing to risk this kind of catastrophic outcome that might be rare this kind of like not catastrophic but bad experience outcome in order to get all of these other gains right clearly there are many cases where the gains are worth it right there are a lot of reasons why people use chatpt and claude and so on to help solve problems today where they wouldn't have before but you have to be very clear about what what the risks are and I think this is where you know legal professionals and doctors and others have been proceeding but proceeding with some sort of caution is the wrong word but let's let's call it with with responsibly in order to balance the risks against the rewards I think in many cases the with things that are consumer oriented you know, you can be a little bit more risk tolerant as long as the person doesn't drop it.

55:11 Whereas there are other cases where if it's really safety critical or you only get one shot, I'd be very careful. >> Thank you, Michael. We probably have time for one more question. is there one that you saw that you know you really really wanted to get to from from the Q&A? M >> what is the the most burning question that you see there? >> Fred asked an interesting one in the chat about whether there are use cases where it's worth building wrappers around AI models to enhance the UX.

55:49 This is going to be Michael's hot take corner. I think there will be use cases. I think many of the startups that are not frontier labs or neolabs right now in this space are effectively rappers around frontier models. This is risky long-term possible short-term. And what I mean is I really do think that good user experience is a differentiator and people will go to things that are more integrated with their workflows and easier to control and understand.

56:32 It does expose you to what in the in the course I refer to as sort of this gravity well of chat GPT which is to say there's an old term about being called you got sherlocked. This used to be when like Apple would take whatever your great feature was, your app that you built for for Mac, and then they would just build it into the OS. And you're starting to see this too where the, you know, the frontier modeling companies, Anthropic, you know, launches, you know, a a design tool, right? Well, Figma make and and lovable and others that build these things themselves, wrapping around them.

57:07 Yeah, it's very easy for the company to then build a UX around it. So I'm I think that the short term is like yes and I actually think that the next 10 years is going to be a story of of experience improvements driving adoption of AI. But I also think that you are exposed in doing so to the risks of getting sherlocked of getting sort of absorbed by the company that built the underlying model. And so the best position to be in is one where you have a modeling mode as well as a product in UX mode. It's much harder, but that's the ideal.

57:53 All right, thank you very much for addressing that final question, Michael. and with that, we're going to wrap up the session for today. I want to take a moment and thank all of you for joining the session today on what AI can and cannot do intelligence augmentation in practice and Michael of course our gratitude to you for for leading this webinar. We hope that the discussion provided valuable insights into where AI creates reliable value, where it falls short and how organizations can use AI to augment human intelligence more effectively. If you'd like to continue exploring these topics, we'd like to remind you that there's a course AI powered product innovation that just that does just that. we will drop the link in the in the chat and you can also get more information by scanning the QR code. if you have questions about the course and would like guidance on where it aligns with your professional goals, our global alumni program advisors would be happy to speak with you and we will also drop a link there so that you can schedule a call and get more information with one of them.

59:00 thank you again for joining us. We look forward to welcome you welcoming you to Stanford online's learning experience. if you'd like to know more about the program, we will drop the link once more, and we will also drop the links to schedule a one-on-one call or register to the program if you are interested in doing so. Now, thank you all again very much for joining and thank you again, Michael, for hosting today. have a wonderful day wherever you are and with that, our webinar is over. Thank you all very much.

© transcribe · For agents Built with care and craft by Gokul Rajaram