transcribe

Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise Diffusion

Sequoia Capital · 1h 5m · transcribed 24d ago
More from Sequoia Capital Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Importance of Being Wired In

Why should college students engage with Twitter for their careers?

Engaging with Twitter and following key accounts can significantly enhance a college student's career prospects by keeping them informed and connected.

  • Being active on Twitter can provide valuable insights and opportunities.
  • Your social media feed can influence your career trajectory.
  • Staying informed is crucial for success in today's fast-paced environment.
# 12:58

Transitioning to the Cloud

What challenges do enterprises face when moving to the cloud?

Enterprises need effective solutions for securely storing, sharing, and managing unstructured data as they transition from on-premises systems to the cloud.

  • The shift to cloud requires new strategies for data management.
  • AI has been explored as a tool for improving enterprise data handling.
  • Understanding the historical context of AI development is essential for leveraging its benefits.
# 25:57

Evaluating AI Models

How does Box evaluate the effectiveness of AI models?

Box conducts extensive evaluations of AI models using domain-specific tests and real-world applications to ensure accuracy and effectiveness.

  • Continuous evaluation is key to improving AI model performance.
  • Domain-specific evaluations help tailor AI solutions to specific industries.
  • Customers can choose models based on their needs from a diverse model garden.
# 38:56

The Role of Systems of Record

What is the significance of systems of record in the age of AI agents?

Systems of record must integrate effectively with AI agents to enhance productivity, rather than creating standalone solutions that complicate workflows.

  • Integration between systems of record and AI agents is crucial for efficiency.
  • Companies must prioritize building superior agents tailored to their specific systems.
  • The competitive landscape requires a focus on both product and agent quality.
# 51:55

Challenges in Knowledge Work Automation

What are the key challenges in automating knowledge work?

Automating knowledge work involves addressing unique workflow shapes, data formats, and change management processes that differ from coding environments.

  • Knowledge work automation requires tailored solutions that consider existing workflows.
  • Data integration and access controls present significant challenges.
  • Understanding the differences between coding and knowledge work is essential for effective automation.

Transcript

0:00 I want to send an emergency alert to like everybody who's like a sophomore or junior in college and just be like follow these 20 accounts on Twitter and and also join Twitter because like this will just help your career. You're either like a year ahead or a year behind simply based on your feed and my feed is like so so wired in. I'll still talk to to 20-year-olds that are like yeah like you know I I see some articles and I'm like what do you mean how do you see articles? Like I don't even know what that means. Like do you just like get lucky that somebody emailed you an article? You just have to be wired in. I I enjoy it. It's a lot of fun. I play with everything.

0:50 Today I'm excited to welcome Aaron Levy, founder and CEO of Box. Box is building a collaboration platform and content management platform for the enterprise that has now taken on really new life with AI. And I'm excited to chat with you today about Box and AI, about your general thoughts on AI because you are such a thought leader and and about how founders can reinvent themselves for this AI wave. So, thanks for joining us. >> Thanks for having me. I'm a big fan of the the podcast. I religiously watch every episode and and so great work. I am wondering though, I know this is a different podcast, but when do you talk about root canals in this? Oh my gosh.

1:23 >> Is that Have we call in Doug Leone? Have we figured out how to weave that in yet or or not? I mean, I'm happy to talk about different kind of pain and suffering that that I've been through. >> We should just do this pod with a root with a live root canal. Let's do it >> and see how it goes. Actually, that would it be like hot ones, but you get a root canal and you have to talk about your strategy while you're in a dentist chair and Doug is just on the other end of it.

1:45 >> Amazing. Amazing. Doug is so pleased with himself, by the way, right now. okay. >> has he seen every tweet? >> Oh, yeah. >> Okay. Oh. Oh, yeah. He's he's very pleased with himself. >> Good. >> okay. Let's start with I'm curious your take on this. Application companies are the hottest Neolabs. Agree or disagree. >> you know, two years ago, I think it would have not made that much sense as like what does that mean? but but very clearly I I think this is what what's playing out in the market and it's all working out mostly because of open source. But what's what's pretty amazing is right now you know the the whole concept of being a a sort of LLM rapper or model rapper is actually working out because what what what I think people underappreciated was that in the real world in the enterprise what you need is some bridge from the models capability to the actual workflow that the enterprise has and that bridge you basically like probably a trillion dollars has been bet on on basically one of of two outcomes. Either that bridge is is very kind of limited or that bridge is actually very vast and and needs to be able to to you know take on you know a lot of depth within organizations. and that bet basically looks like are you only long you know kind of the the model itself and super intelligence or are you long this sort of application tier or you know maybe previously would have been just pure neolab but but I think it's very clearly playing out that that actually there's a lot of gap between the model and the workflow and as you bridge that gap over time you get to a point where you realize oh I should also do the model and and then you have enough data you have enough sort of domain expertise where that becomes its own you know sort of flywheel So I'm very bullish on this on this idea and what's cool is it's opening up like multiple layers of of you know sort of opportunity for startups because you could either be the actual applied company itself i.e. the Neolab or you could be the infrastructure provider to the Neol and you have like multiple layers of of you know going and attacking that space but I think a huge update for for the market and everybody's kind of view of it.

3:47 >> Box Labs let's go. >> We we already have it. It's a little bit it's a little bit secret but we're we we we pay very close attention right now. the main focus is let's make the agent really really good on any model but over time obviously you would peel off certain use cases either on a per customer basis or or kind of across the the whole data set >> yeah I mean it seems like there's two forces that are happening one is people don't want the fox to be guarding the hen house they don't want the seller of the token to be the one that is also metering and gating you know what is the best token for each use case yeah >> and then the second is you know there's actually a lot of work to do on that bridge to cross and there's real research involved in it And it's pretty bespoke to the exact the exact workflow and the exact end customer that you have.

4:29 >> Yeah. I think I think what what I tend to see happen in the valley is and and for very good reason and it's actually why these companies have been so successful is like everybody is so kind of research pilled which is again totally awesome big fan. It's it's led leads to all these breakthroughs. but the the sort of sense that okay the model and the intelligence in the model is kind of the only form factor that matters. And then you go to the real world and you you sort of see how intelligence actually gets rolled out in people's workflows and the model could be the most intelligent you know super intelligence in the world but that workflow still requires you to connect up to other data systems still requires these moments where there's a human in the loop interaction. There's delays in the workflow and so the the sort of agent has to sit idle. there's change management of the actual business process. There's legacy systems that haven't been modernized. So you kind of go through these five or 10 things that are way sort of, you know, much more operational, much more blocking and tackling than just the pure super intelligence of the model. And the last thing that I think a classic sort of research organization wants to go do is go attack every single one of those things. And this is sort of not a there's like this is not a sort of a one-off in history. Like we've always had this relationship between infrastructure and application. like you know obviously AWS or GCP or Azure have created you know trillions of dollars in value of of kind of market cap of infrastructure but guess what there's also trillions of dollars of value in software that only exist because of that infrastructure and if you were to go back >> you know 10 years ago and you were to look at what was happening in the data space and you looked at what what you know GCP early kind of versions of GCP was building or or AWS is building I guarantee you would not have predicted Snowflake or data bricks existing you would have been like the infrastructure just already does that why would you pay another $10 billion of revenue to all these other products that are just making it so you can work with your data. The same thing is going to be true for intelligence which is the models will be insanely valuable but the application of bringing those models into real workflows in banking and life sciences and healthcare and government that's just going to be a lot of software. Now the sort of challenge for the next let's say two to five years is how much do the model providers need to move up that stack and also try and you know go and compete at that layer or do they either leave open intentionally or accidentally that entire space for the application sort of ecosystem and actually to some extent this is like a a big question strategically for them because on one hand you want to be closer to the customer on the other hand you also want to be able to have an ecosystem so people trust you so so that's going to be like a really interesting tension over the coming years.

6:56 >> How do you think that'll play out? I feel like that's the question we spend every day wrestling with. >> Yes. you can Yeah. You know, I don't want to speak for for Sequoia, but you can sense some of the existential dread of a VC right now is like, should I just put another billion dollars into Enthropic or, you know, should I attempt to sort of see what what at the applied layer is going to is going to play out.

7:15 So, nobody's, you know, that envious of of your guys position having to to figure that out. But but obviously, it's even harder for the entrepreneur. but you know I I am you know I'm pretty long obviously the application layer like I'm also equally very biased like I'm I'm I have a very concentrated bet u with very limited diversification on on it working out that you still want to buy technology that that sort of understands your workflow and can get to the the core enterprise data. But I don't know that that there's I I just don't see a different event happening than the than all of history. I mean, you know, we only have like 50 or 60 years of of of computer software history. but basically, when you go to that law firm and you go to that, you know, pharma company and you go to that bank, they need something that bridges the core technology to their workflow in their business process. And AI has not meaningfully sort of changed the need or the shape of what that looks like. And the and the best sort of manifestation of this is sort of this chatbot versus kind of agentic workflow, you know, demarcation. So the chatbot can be totally universal and can be totally horizontal, but you know, all of a sudden you kind of look at that and you say, "Well, my workflow kind of needs something to kind of ping me at the right time in the process or it needs access to a certain kind of data that that just, you know, the the chatbot can't natively get access to. So somebody has to go into that organization like get it set up and somebody has to go and provide domain expertise to this model so it really understands our particular business process." and and so then unless you really just underwrite, you know, the big labs at at honestly like I'm not exaggerating, like a 100,000 employees, if you don't underwrite that, then then the diffusion kind of economy is going to be massive because every single one of those companies, whether that's a 50 person firm or certainly a multiund,000 person firm, is going to need an army of people to go in and help them with that transformation, that change management.

9:09 And so this is, you know, like whether it's the FTE phenomenon or just again understanding that domain expertise that that's going to be a very big deal. And then you you kind of alluded to to this point Fox kind of guarding the hen house and and that that might actually be singularly the biggest sort of reason this has to happen which is even even under all sort of like complete benevolence and like nobody's actually doing something in any kind of like you know sneaky way it just stands to reason that if I'm going to sort of give a task to an to an a an agentic system I just want that task to be cost optimized like with accuracy as as kind of holding constant And so who can do that? It's it's the company that doesn't care among you know 10 different models which model is is performing that task. So like definitionally you would want that to be the the the the company that does not have a sort of a preference.

9:58 >> And then the and then the force you have going against that is that the model companies can subsidize their their models or offer them at different different rates. >> I don't know how long that lasts though because when when these companies become public, I think they will be held to basically the same laws of of capitalism that everybody else is. So, so the subsidization is working very well up to a certain threshold of spend and we might have we may have exceeded that spend when you're at like tens of billions of dollars of of kind of capital and in fact if anything it might even be worse because eventually you have to go in and sort of pay for your training runs as well. So like they they like the subsidization of tokens is a I I think it just has to be a temporary phenomenon.

10:37 >> You mean from a gross margin perspective or from like an antitrust perspective? >> Entirely gross margin. Like if I >> but they have such high gross margins on inference right now. >> But then what are they subsidizing the like then then it's actually then they're then they're charging actually like you know decent rates. >> Yeah. But they can afford to price the API higher than their own first party products. Right. >> So that part to totally fair. You know there's an interesting kind of you know dimension which is well like if API is the high margin thing that is sort of paying for the subization but you've moved all your customers over to the applied product and there's no API revenue. So like there is like an equilibrium you have to strike with this.

11:11 >> Totally. and and so then all the while if you have some sort of either non-economic actors or people just with a totally different game theory in this you know Meta being one maybe SpaceX being one China certainly being a giant one even Nvidia being one like that changes the calculus as well which is those four kind of cohorts don't necessarily need to make money on on inference in the in in sort of at least at the same margin structure at like what anthropic or openi needs. So they would might be fine to bring down inference to 10% margin because they just want to basically you know pay for massive compute clusters and then as long as that happens and as long as like there's not insane proprietary sort of you know you know sort of closely held secrets then no matter what you're going to have kind of cost per token go down on a on a like forlike basis all of which means more value acrru to the application layer which is I think like a very long-winded way of saying I think there's just value in kind of everybody in the stack. I just don't know that that the I I the only thing I probably wouldn't bet on is just okay, one or two labs get 95% of the value creation. I think there's just going to be a much more dynamic environment. And honestly, if I were one of the, you know, two or three biggest labs, I think I'd prefer this outcome too because back to your antitrust point, like at some point you'll just be nationalized if you're the only thing that that is is sort of exists as intelligence. So, you kind of want a little bit of healthy competition in this ecosystem anyway.

12:34 >> Yep. Totally. let's let's transition talking about bucks. I'm going to come back to talking about tokens and Jeban's paradox and China and all these all this stuff, but let's let's talk about Box for a second. >> I love that topic, too. >> give folks I mean I'm guessing a lot of people that listen to this podcast use Box, but give give people like a brief kind of explanation of the history of Box and how how you're reinventing yourself with AI.

12:54 >> We started the company as a way to be able to kind of securely store and share data in the cloud. and and it was a very simple idea, but but we we we kind of just kind of cracked a a nut or struck a nerve. and that you know for for Doug out there we we were able to sort of scale up quickly. We pivoted rapidly in the enterprise and the idea was enterprises would be moving from on- premises systems to the cloud and they would need a a better way to be able to sort of store, share, collaborate, manage their all of this unstructured data, their their corporate documents, their financial documents, their marketing assets, their research materials in the cloud securely. So that was the the kind of company and we we had been sort of flirting with AI kind of products and and sort of experiences really since like 2015. if you remember like the first kind of rise of the you know maybe the the at least in modern times the the AI winter that happened in like the 2015 to 2018 period which is like we think it's going to happen now and then it didn't and that was a period where we were like okay these very you know sort of early AI models were showing us signs that okay if you if you you know looked at an image and you could classify the image well that's pretty useful if you're in an enterprise because now maybe you take all of your image data and sort of label it or maybe you'd OCR something and you' be able to kind of pull out, you know, kind of, you know, sort of the the the text in there. That's enormously helpful. The problem was is insanely expensive and you had to have a model for every single use case required, you know, kind of a hyper trained model for each workflow that you wanted to do. So, we kind of shelved it. you know, few years later, you know, started paying attention to the GPTs. We had some hackathons where people were like, "Oh, we could do like, you know, type ahead in one of our kind of, you know, note-taking products." And, and that was that was sort of early kind of versions.

14:40 We did some some early work in in sort of text detection and classification which helped with security use cases. Then obviously catchb moment sort of hits and and that was the the big sort of you know head exploding moment. If for no other reason why you know then then they kind of figured out a form factor that opened up everybody's mind to oh these could be these interactive systems that you just like ask a question get an answer back of of sort of you know kind of increasing complexity and length. Yep.

15:08 >> So we looked at that very quickly. We jumped all in. We sort of did the whole company pivot like like everything was was kind of exactly like academically what you should do. We had a team carved out. We we put you know the best people on the team. We we you know met every day looked at the updates and then slowly but surely sort of built out what what today is is kind of our you know our AI stack and then basically the box agent and and for us you can imagine the use case is is you know very straightforward. We sit on hundreds of billions of files. Every single one of those files contains critical information for an enterprise. That could be their contracts, their research files, their marketing assets, their loan documents, like all of this critical information. The problem is they rarely know what's actually inside of it. So they they they unless you literally look at the document and and kind of search and find it, you just don't know what's inside of it. So now agents can go and basically be farmed out to go and answer questions about that data. they can pre-process it and extract metadata from those documents and turn it into structured data. You can use agents to automate sort of steps and workflows. So, we've built a platform that basically lets you deploy agents against all of that unstructured data. And that's been the the kind of core focus.

16:17 >> Are you using agents to create new content? >> We are. there's a couple modalities where that shows up. One is we have again an online sort of collaborative product that that an agent can just like generate any amount of content in it. And then and then we've done most of the I think probably more exciting work with with OpenAI and Enthropic on just how do you do like advanced document creation, PowerPoint creation. We've we've decided that their tech is is you know at this point can always be frontier. So so we have an agent that goes and interacts with those systems to produce you know a high quality PowerPoint etc.

16:50 >> Awesome. Your tagline your business lives in content unleash it with AI. So what are the hero home run use cases for how people are unleashing it today? Yeah. >> And then if you had to fast forward a few years, what do you think people will be doing with AI in your products in a few years? >> Yeah. So the the probably the easiest hero for again more of a traditional enterprise to think about is is just as simple as you have a million contracts.

17:12 Why don't you find out what's inside them? Or you have a million research documents, be able to go and pull out all of the critical structured data, put that into a database, and then be able to query, analyze, automate workflows around that. So that that's kind of the thing that that just knocks it out of the park every single time because it's been a long-standing problem that people have never been able to go and and and sort of apply human, you know, sort of labor to because it's just too expensive to read every every contract, every research document. You know, maybe you could do it if you had like a loan document process, but most other data just never gets read at that at that scale. And then I think the the stuff that that we're probably you know as much if not more excited by is is really the equivalent of what we see with let's say coding agents or other other other complex agents which is you have longunning agents that are just executing your entire kind of workflow or process. And and this would be in the form of you go to a bank and you're onboarding at a bank and they've basically like automated every step that is possible to automate and then sort of jumps out to a person in the in the steps in the process for extra review or extra verification. But now instead of that sort of one or two week back and forth, it just happens in like an hour.

18:20 Like that's the dream state of most of these enterprise workflows is what if we could onboard a client faster? What if we could discover kind of critical data inside of our research much more quickly? What if we can alert to a security event much more much more quickly? So to do that, you need these sort of background agents or workflows that are that are sort of pre-established for those processes. >> Totally. It's the year of the long running a agent.

18:42 >> It is. It is. >> Yeah. Yeah. I'm curious like you made the analogy to cloud code. It seems to me that in the coding domain, using AI is like not only accepted, it's embraced. >> Yes. >> in the content domain, which is I think a where a lot of the content in in box sits. Yep. using AI to produce content at least. It's just like there's this almost this allergic reaction to it like all the panggram stuff on Twitter.

19:04 There's the you know it's like this concept of work slop. I'm curious what you think about work slop and like will this still be a thing in a few years? >> I I'm going to sort of separate the box corporate hat and just now kind of maybe riff as a as a as a as a consumer of of >> gosh I wish there was a better term but work slop as you know inside of an enterprise context.

19:26 >> I get board decks that are entirely written by AI these days and it kills me. >> So here here's the difference I think on the acceptability. so there's probably like more symbolism to this actually topic than than than just like the this the slop element, but like actually like diffusion of AI in general sort of almost ties to this code like other than you know the the top engineers that we hang out with that like are like they have you know deep taste in the code like like you know and and the judgment is incredible and and like it is them as much an art as is a science. So take that group aside. For most of the world, code is a utility. Yeah. It is just trying to accomplish something. You you're just trying to automate something. You're just trying to put a a sort of interface up there that somebody presses a button and moves to the next step. So for most of the world, the value creation of code has been to automate things and to be able to have it as a utility. So, so at the end of the day, like like we're and we'll probably still use the term slop for a while because because there's taste in kind of front-end design and there's taste in in sort of systems and you don't want to have vulnerabilities in your code. So, that's going to exist for for a while. But, at the end of the day, if you can tell an agent like, please go and generate my entire backend system or my front-end system, like it's just it's not only acceptable, it's it's preferable because it's just like that was the thing that was blocking us from moving forward. So, we need to go do that.

20:52 At least the way society functions and the way the world works and our brains work at the moment. Maybe this changes. You know, when you when you get a presentation from somebody, there's still this association which is like I'm trying to decide if I can trust that person to go and execute on that thing or deliver that result or understand that topic. And so if when you see work slop you're like I my I'm losing my ability to to sort of know for a fact that like like how much of the of the thought process was them versus how much was the AI? How much should I even care about that because I myself am doing the same thing. So like we have this weird like it's a it's this very weird sort of like collective issue that we have which is like which is like I'm doing work slop for some of my you know brainstorms and decisions but when I get it from somebody else I'm like hm should I trust you? and and I don't know I you know I mean it just might be a thing that as a society we have to kind of keep cranking through over the next kind of three to five years and and end up at the other at the other side like like you know I hate to use like these like totally busted analogies but you know obviously you don't care when you see somebody's financial model you're like yeah that was generated clearly by like a macro or or you know that was like not you did not personally go and compute all of that but but you're showing it to me and we're talking about it So, why can't the same exist for a strategy deck or or whatnot? But I think right now we're going through this evolution of like what is the person's role? What is the content a proxy for?

22:18 Is it supposed to be a proxy for how much that person knows? Is it a proxy for what we think that they can go and execute on? I think I think we're just in this very messy period where we have to kind of figure that out. >> Totally. did he read the Stan Ducken Miller Wall Street Journal? I I read the I read the discussion about it. >> Yeah. The the the reaction to it, but I actually I I didn't I didn't read it. But was it like very sloppy?

22:40 >> I didn't think it was love. I loved it. And so to me, it was just a nice counter example of I have this like allergic reaction. >> How many how many it's not exits wise were in there? >> I don't think there were but it but it does show up as 100% AI in in Pangram. >> Okay. And dashes >> there. I think they're okay. You can't do that. Yeah. Yeah. But it's like it it was a nice counter example to me because normally I read something that's clearly written by AI and I just have this allergic reaction whereas with the Stan piece I didn't. Yes. And I'm not sure how much of that was just you know it's Stan therefore I trust in Stan versus >> Yeah. No but it's it is psychologically kind of like weird because I I'll read these like you know X articles and I'm like now like doing 2x the amount of work to read these. I'm reading it one for the substance and I'm also reading it two for the calculation of like did the person write it or am I just literally reading like a claude prompt and then like I'm like I'm literally like my mental processing is now like should I now does that upweight or low or like lower my my sort of judgment of the person or the post and I think we're yeah we're in for some weird times because of this. I would hate to be a college professor. I would just I would totally quit because you're just like you're like I don't I don't know anymore what you did. I I like what what's >> it does seem like the calculator is the closest analogy though.

23:50 >> Yeah, it except it's just like that was like more like finite in terms and you still had to piece together so many more things >> like we like the and some of these analogies are breaking down of like it's just a task because it's like well at some point like this thing is doing like at least 10 tasks at once. Yeah. >> but but yeah, >> let's talk about harnesses. How does the how does the box >> Great transition to harnesses.

24:12 Okay, >> speaking of calculators, let's talk about harnesses. Yeah, >> this is back to the kind of Neolab kind of applied applied layer. there's there's a bunch of things that that we know about our file system, permission structures, our search our our search engine that that you know certainly and by all means we actually would love the the we'd love all the labs to train on our understanding of this because it would only make external agents you know as good as possible. Yeah. like like we we we always talk to labs like hey we'll give you as much data as you want about about kind of how the system works but but let's just say like that aside we we have a lot of depth of understanding of what do people do in box how do they search box how do they decide when they look through 10 files which is the one to go pick how do what what is the in their internal kind of calculus or heristic on on sort of figuring out the most relevant document to look at. So we know all of that and we basically you know have have built an agentic harness that that you know attempts to kind of understand that set of domain understanding about our system. it it obviously has access to our search system our file system. It has a a bunch of kind of mechanisms for just pulling out just the text of a document just pulling out chunks from the document doing embeddings on the document on a fly. So there's a set of kind of tools that it can use. and effectively it's a harness for asking questions of a large data set. So in my box account I have I don't even know the latest number but on the order of tens of millions of files. just because it's like it's like every everything that has ever kind of accumulated over over 20 years. But I can now ask any question of of all that data set using the box agent and it it goes around it does it does a multiple searches in one. It reranks it. It then very quickly sort of pulls out the most relevant information. And then it'll in some cases read the full document, you know, does all the steps >> and and then we compare that against like, well, what if we just gave, you know, claude our API or gave OpenAI our API and we see like meaningfully better results on accuracy and latency because again, we kind of know exactly how to how to tune it for our workflows. so that that's effectively the the harness that we built out.

26:15 >> And then what EBLs matter the most to you? >> so I have a couple I have a couple like, you know, just funny personal ones that I just keep track of of like my own use cases. But we do we have I don't know hundreds of different tests that we do on every single model. we actually have two evals at the moment. One is we put out a thing called the complex work eval which is a which is a set of domain specific work in life sciences financial services public sector tech etc. And it's it's kind of exactly what you'd think of as a as a document ccentric eval. So given these five documents and this set of problems like what would your answers be? And we test every single model against those with our agent. And then we have a hold back eval which is actually the first one is hold back also but but the second one is just like then our box instance and how box employees use use their data and then we eval every model again on that.

27:03 so we're able to kind of roughly keep track of of of all the incremental progress like we we see when things move by half a point in terms of model improvement. and then we roll out sort of default models based on different kind of cost and accur accuracy thresholds. And then we let customers also choose any model they want from effectively our model garden. >> What's your current view of the race and like where all the horses are in terms of model performance on your use case?

27:28 >> They more or less closely correlate code with one exception which is actually in some of our use cases Gemini is disproportionately better than than what you would see from coding. and and it might be, you know, sort of just better tool use. you know, given given kind of, you know, the the the Gemini ecosystem and what they need to build for it it it solves, you know, kind of a a strong set of sort of general knowledge work use cases as well.

27:55 but I think mostly correlating to to kind of code. So, Fable 5.1 was clearly kind of, you know, state-of-the-art and the best model that that we've seen. There's obviously rumors about other models and so we'll see how the kind of race y you know kind of you know continues on this on this front but but basically by and large like when you look at GDP vow Merkore has their Apex eval these things will all generally follow the the coding models and and so I think it's we're just neck and neck on like gro muse the fable class and you know GP56 whatever they're building next like these are just It's a total race right now.

28:35 >> And do your customers typically express a preference on which model they want to use or do they just use your default? >> So they they you know kind of by volume they use our default because it's just easy and and it works extremely well and it's tuned for there's a a few ways our our kind of agent manifest. So like the way that you'd most commonly experience it as an end user is you would just be searching and and kind of you know asking questions of your data.

28:59 the but by volume the the volume of tokens tends to go through more of our workflow agents or data extraction. That's where actually you you have customers actually doing eval basically saying okay I want to really make sure that at this cost profile I can get 98% accuracy on data extraction and that's a place where like we'll have an FTE that goes in and helps you you know helps understand your data environment tests against five different models and then you're just basically at the mercy of of the eval.

29:25 >> Yeah. What are you seeing in terms of the adoption of openw weight models in your customer base? >> so probably higher than people think, lower than what enterprises actually want and much much much much much lower than what it'll be in five years. So, so like some mix of that would be the kind of message like >> and is primarily cost that's driving that decision? I I you know I have to probably attribute 30 plus percent to just kind of the the sexiness of of of like >> I want to try GLM.

29:55 >> Yeah. Yeah. I think there I think there's that like I've heard >> I've heard you know CIOS of Fortune 500 companies say we're we're playing with open source here and I look at that and be like well I know for a fact like >> like Gemini or Muse would have been just fine at that particular cost profile that you're trying to do or probably even like you know you know 56 Luna or Terror or whatever whichever you know one had the had the crazy discounting they just did like it probably would have been totally fine but you want to be able to be like okay I'm a little hedged like it's cool like like we're at that phase still over time I think it stands to reason that that you'll see meaningful different costs because you'll be able to peel off workloads that that just only make sense at at sort of you know grinding down to the cost of inference in which case open open weights will have the sort of economic advantage. Right now there's you know this challenge of like sometimes it's more token inefficient you know sometimes like randomly like I've heard stories like randomly it'll just like speak Chinese like like mid midchain so you're like okay well that'll be weird for a bank. so funny. So like we need to like probably work on some of those things. But longterm I think it has to be, you know, the case that that you're going to you're going to peel off those workloads. You know, one one of the more interesting posts I think on this that I totally subscribe to is Jesse at at Decagon. Do you probably read that post of like of like this paradox of like we're going to have like you're going to see clothes just, you know, continue, you know, go exponential. But what's going to happen is each use case that kind of matures you can peel off to open source. And once you have kind of stability in that use case, it starts to make sense to veer it toward an open weights model, assuming one of two things is true. One, that that it's actually literally cheaper, or two, having some post- training gets you x x% more performance. And and so I think you will just be in a reality where we will we will and this is going to be very confusing probably for like the press more than people in the valley because you'll be like, wait a second, like the revenue of entropic, opening eye, etc.

31:44 are like off the charts but somehow open weights is like also growing exponentially and you're like how is this like how is open routing so fast? >> Yeah. And it's like the pie is growing so fast >> but what's happening is is actually there's there's an interesting duality. It's it's it's not even just like like rising tide lifts all boats. It's like no no like we either use Fable or or 56 for orchestration and then we farm out all these longtail tasks to a cheaper model. or the opposite is true.

32:11 like you use you have some kind of orchestration agent that like by default does the cheaper stuff but occasionally sort of sees some something that is just way too hard and then it pops it out to to one of these heavier models. And so you might have blended 50% spend on each but 10 times the amount of tokens, you know, on the open weights model. And so like everybody's kind of winning, but there's an there's an interplay between why they're winning. I'm curious about how you think about memory and customization or personalization and where that's going to go because it seems like today the dominant architecture is kind of like rag based system still like you get fancy on the rag but it's still context lookup where the weights themselves aren't fundamentally changing. Yeah, >> it seems like I mean if I listen to my friends at the labs like continual learning this idea of like the models weights should adapt as it gets to know you.

32:58 >> Yeah. >> we had engram on the podcast. I don't know if you if you know Dan like >> I I just got introduced to him. I listen to the podcast. Soing I would have loved to been the fourth person in the room or >> Amazing. Yeah, they they I I think they're working with customers that to to help basically bake in some of the context into the weights themselves. What direction do you think this going to go?

33:18 >> You know, you're catching me at a time right before I'm actually doing my call with Dan. So So I I wish I I could have talked to him first and then I'll have like a way more eloquent answer. I'm like extremely fascinated by by the approach. like I I have no reason for not wanting to it to work and and exist. We live in a world at Box where we see these sort of the the the high degree of complexity on permissions and access controls and data that that tends to be sort of the rub on a lot of these types of of approaches. And I'm I'm going to put engram aside because like I'm sure they've already thought this through. So I'm I'm going to talk more generic phil philosophically. I think sometimes you will talk to a researcher that you know sort of imagines the world working the way they work which is like I'm a researcher I have access to everything and so if I had a model that was trained just on my world this would be amazing and then you're like let me introduce you to a lawyer and the lawyer like has this tiny little you know access point of just like the five projects they're working on because somebody right you know one door over is working on the competitive project to to a another company in the space that they can't have any sort of overlap with with what they see or what they know and there can't be a single document that passes between those two walls and they have to be these kind of hard barriers. So, you know, so like sure, like you could still you could still now train a model just for that one user, but like what happens if every single day they get added or removed from something that that adds important context to to sort of what they need to understand. And and and you know, again, like I think there's going to be probably breakthroughs in continual learning that that sort of all resolve this. But like this is why previously it was just like, you know, there was no way you could pull this off 5 years ago because it would be insanely expensive, impossible to kind of wrap your head around around how those access controls are supposed to work. But, you know, obviously with like as the cost curve goes down, as open weights, you know, get, you know, cheaper, smaller, faster, better, I think this becomes super interesting. One thing on on the podcast that I found very fascinating and I just need like a T-chart honestly is just like, you know, like what is the decision point of what goes in context and what goes in the weights? Like you you have to be a little bit thoughtful about like where is the the massive performance gain that you get by by baking in the weights. And there's probably some like incredible like like calculation of like like when the rate of change of the data is not you know so far but the upside of the of the weights you know you know dramatically change the accuracy of the model like you know you'd have to kind of land on some sort of you know rubric like that.

35:49 I mean, if you could wave a magic wand, it almost seems like I think Karpathi has said this in some prior interviews, like if you could almost remove all the memorized information from the models and just have have it encapsulate the specific >> reasoning capabilities like the ethos of how we do things for example at Sequoa and then you have all the you know actual content and a lookup system that almost feels like the if you could wave a magic wand that's what the system would look like >> and so so that one's super interesting and the I I think the the the question will be like how much are enterp enterprise is different at that level versus it's actually their their literal IP that is what makes them different like how many different types of styles of execution are there in the world versus no it's just like the depth of knowledge about that particular legal case and and how do I apply it to this other project I'm working on that's where so much of the value sits so so but again if you can just like wait till my Zoom call with Dan and then I'll really know the answer but like I'm a fan because no matter what there's going to be like like I've jumped right into like the individual, you know, but like no matter what, like at a firm level, there's probably ways to take this approach. Like I'm a big sort of fan of what Trajectory or Applied Computer doing or Prime Intellect because because I think there's like there's there's no question that if you're Eli Liy, you want a model for how you do drug discovery and that probably does need to go like farther or deeper or more sort of specific than what you're getting at off the shelf. And there's not a lot of sort of, you know, kind of church and state problems for for drug discovery, you know, workflows. they they probably want as much of that information available to as many people as possible.

37:19 So I think it's going to be like domain specific and you know you're just going to have different outcomes based on which vertical or or you know type of use case and where the firewalls need to be in that process. >> Makes sense. Okay. So you hinted at the beginning that there's a Box Labs. What what type of work is Box Labs doing? >> so it's the equivalent of Box Labs like I don't know if we've used capital L yet. but but basically you know it's our it's our applied sort of AI team. And >> what research areas are most interesting to your your team right now?

37:45 >> Yeah. So, so the the the biggest areas that that the most sort of research kind of or you know of the continuum of engineers there's you know some cluster that is sort of more on the research bent and of that cluster the things that we spend time on still again is at the kind of applied layer but but it's a lot around how do you how do you take agents and make them you know another 10 points of accuracy improvement given x problem. so what what is the you know how do you build a map of the of the problem set you know with a given set of data to sort of best execute on that task. So we spend a lot of time on on that style of work. We have a team for instance working on how do you do effectively at the a agent level at the harness level you know some form of kind of auto research on on sort of hill climbing on accuracy of answering questions or set sets of problems on on given a kind of a set of client data. So if you're a bank, you have a bunch of loan documents coming in and these are like 100page documents like whether you're getting 70% accuracy with an off-the-shelf model or like 97% is like basically obviously a world of difference in can you actually go and automate that process. So somehow you have to hill climb from the base model to the 97% and there's a lot of work going into the system to to basically pull that off. Maybe zooming out, what do you think of as, and this can be a box question or a non-box specific question, the role of systems of record in a world with agents? And I'm sure you saw some of the Twitter discourse on like, you know, every software company is trying to sell me their own agent right now. I don't want another agent from them. I want I want their system record to work well with my agent. And so, like, how do you think about that?

39:21 >> Hashtagcloud force. so, >> catchy name, by the way, very catchy. >> I I mean, literally, it's like one of those things where where like the first three minutes you're like, man, that seems funny. And then four minutes later, especially when you see the stock, you're like, " brilliant move." Like like this is great. We're doing this. and then when you I think somebody somebody actually said this the best. They were like when they heard Matthew McConna say it out loud, that was like really the that sealed the deal. That >> was the aha moment.

39:47 >> That was the aha moment. I was like, man, he can sell software. Like it's actually incredible. Like his voice is so good for selling systems of record and agents. if you're in like like our sort of you know contemporary group of like you built a SAS platform and you have some set of data and workflow that that you know your customers kind of operate in there's there's effectively two things you just have to do and I I think anybody attempting to do one over the other is just going to lose. you have to build an agent that is insanely great at your product. like that agent has to be you have to provably be like 10 or 20 points better than an off-the-shelf agent at at using your system. Not because it's hobbled the other side.

40:31 It's just like you are so Eval maxed and you're like so tuned to your particular workflow that that you can improve your system. You have to have that and you probably because you understand your domain unless you're like totally asleep at the wheel. you probably have use cases that no one has has sort of thought to build products around because you you talk to customers every day and you see what they run into and you're like, "Oh, we could just like have our agent go do that for you." Like I've had at least a dozen I mean, so I probably talked to a couple hundred customers a year in in variety capacities. I've had at least two dozen times where the customer has a use case that is like a breakthrough moment for me of like, that would be actually totally insane. like what's example?

41:12 >> unfortunately, since you put me on the spot, I don't know if my example will pay off the the level of excitement I just had because it'll be >> put you on the spot. >> No, no. I It's just like I think like the thing I was thinking of is just like I think it's going to be a wamp wamp for the podcast. But there was there was basically a customer had this idea of they wanted an agent in the background trying to sort of figure out when documents sort of met their governance policies and like does something need to go into some kind of archive or something into some kind of legal hold or whatnot. Oh, cool.

41:43 >> And and see exactly see that's exactly the voice that was exactly the voice I was worried about. Yes. Yeah. No, you couldn't even you couldn't even pull it off. Okay. So, so but in our world, this is awesome because you're like, "Oh, >> yeah." >> Like, no, because think about it. Every company has a head of governance. >> I meant Cole sincerely. >> Okay. I know. I believe you. I mean, listen, you do enterprise, so like I think that it was at least half serious.

42:05 imagine you're an enterprise. You have a head of compliance and a head of governance. Okay? They can only be like overseeing the whole sort of enterprise. They've never been able to be everywhere at once. Now imagine if they could sit next to the employee and and basically be able to be like, "Oh, you're about to go do something that breaks our governance policy." So So like the idea was like, "Oh, what if there was just an ongoing agent that just like automatically was just like, nah, that's going to break your governance policy."

42:34 Instead of the user having to like try and predict or understand this stuff. So anyway, those are the kind of things where if you have an agent within your product, you're going to be able to identify sort of sooner and better than than the rest of the market and or just like do things that maybe would be impossible to do off platform. On the other hand, it's just like obviously you have to go headless. You literally have to make sure that your APIs are exposed to cloud and chatbt and and you know all the different platforms. And you have to make sure that you have a either a direct way into deterministic APIs so that those agents can use your your APIs to make calls via MCP or whatever or at least make your agent be headless and be exposed in those systems. And then the only reason maybe this is like a like a even remotely a hard debate is you have to just make sure as a as a system of record that you can find a way where commercially it sort of makes sense on the other side and and is sort of valuable and interesting. And the reason why I think a lot of people got that wrong that that weren't in these companies was just underestimating the amount of of sort of new use cases that are just total upside. They're like complete whites space opportunities for these systems of record. So like in the Salesforce example, I use Salesforce more today, probably by an order of magnitude than I ever have because I MCP into it via Claude or Chatbt and so and so I just am always asking questions about the data inside of our CRM system.

43:51 >> And do you think that means the systems of record become more toll booth businesses then to make sure that they're capturing the opportunity? >> I don't love that term. because like no one's had a good experience at a toll booth. >> I love toll booth. >> Yeah, exactly. You love it more than governance agents. so so so I would say that that because they they sort of have a have a depth of purpose of organizing the workflow, managing the data, securing the data, providing guard rails, then it's really just yeah, you need to you have to have some kind of volumeoriented business model on that other side. And and I just think there's like if you're solving real problems for customers, it'll just like make money.

44:26 I've like like this is so cheesy but like I've I've told like LinkedIn product managers I'd probably pay 10x more for LinkedIn if I just could MCP into it. If I just had a way of just like always understanding like like okay this CIO is doing this thing and I need to reach out or whatever I'll I'll like take my money. Yeah. >> So so there's these systems actually have a tremendous amount of value based on the data that they that they have and customers will absolutely sort of find some way to reward you for that value creation if you're doing a good job.

44:55 >> Super interesting. maybe related let's talk about product UI. >> Yep. >> Big you know generic chatbot chat box agent. is that going to be the dominant UI for how people are using AI in the future especially when it comes to the application layer? This is where why I think the applied layer, you know, has so much run u room to run is because probably the the sort of, you know, universal chat system that you ask a question to, you get an answer back or it does sort of some work in the background. I think it's going to be obviously like that's going to be a mainstay that that that UI will always exist. It'll be incredibly powerful. The horizontal products will have it. The vertical products will have it.

45:35 Everybody will have it. It's just like >> your product has a search box. Obviously, it does. Yeah. So, so that that's always going to be here for these sort of one-off asks of an agent or go find this thing or answer this question or produce something for me on demand. But most of the enterprise is sort of made up of these processes and workflows that are kind of just happening behind the scenes. Sometimes they're happening with computers and computers are running these things or sometimes they're happening with other people that are doing these things or sometimes they should be happening with people but you could never afford to have them happen with people so they just didn't happen.

46:06 And so that's that sort of is a slightly different kind of, you know, metaphor than than a chatbot where you ask a question and it comes back w with an answer. That's like, okay, I I kind of want agents in the background to do things for me, read every contract, look at every log, triage every security incident in incident, and then instead of me chatting, maybe I'll chat as a as a means of doing kind of a catch-up. I want a dashboard. I want a workflow. I want a queue. I want a task list. So then the the the the challenge becomes well does the horizontal product kind of take on every one of those those components and and manifest every one of those experiences in one in which case I think you'll start to be like man that thing is like really like a that's pretty heavy and like and then we'll start to like be like oh this is no longer this simple easy delightful thing anymore. So then the vertical players actually like like they actually sort of understand the process and can manifest all the right buttons and tabs and the names of the things for that particular workflow. So I think it puts I think as you have agents that are doing more work in the background doing more you know async work that is just like I farmed out a bunch of agents to review things as they happen or whatnot that leans more toward the applied companies that can understand those workflows that can understand those processes. I think you're going to have these in every field. you're going to certainly we we already know that you know how they're going to look in legal with Harvey Lor etc. we are seeing them start to emerge in areas like security. We've seen them start to emerge in the longunning kind of coding agents with cognition and factory. so I think that that will be you know one of the bigger kind of applied AI sort of use cases and ultimately like in five years from now I would bet like 90% of all tokens in the enterprise are things that a user never kicked off and they just see a result. They just they see a task show up and they have to go review it and it's just like it's just happening.

48:01 >> Yeah. Yeah. Makes sense. Okay. Let's talk about AI diffusion. coding agents it was like you know boom January 1st 2026 happened and like the fastest diffusion of anything into the economy we've ever seen has happened. >> the diffusion of you know the rest of the AI magic into the rest of our jobs seems like it's been a lot slower. >> Yeah. >> What are your thoughts on that? And where are the areas where you think we're going to see faster diffusion and how is that going to happen?

48:26 >> Yeah, so you always have to kind of compare and contrast coding versus everything else to really understand the the dissimilarity. so in coding essenti and this is back to the kind of utility point on on you know slop like the utility of code is almost 100% represented by the amount of text that you can generate like like all obviously an insane amount of value went into the text but but like and like knowledge and expertise and meetings and everything but like ultimately like the text is the thing that produces the the program that is actually the thing that you're trying to do. You know, if you could have the world's greatest programmer like and never had to sleep, never had to eat, they they could intuit it what to build and they could just sit on a computer all day long, your value creation would be 100% correlated with how many hours they could sit at that computer and type code. Like lines of code, ideally, ideally good code is the most thing that will be correlated to whether you produced software that that people wanted. So it's all text. The models are hyperrained on these. Everybody in AI labs treat coding as as a competitive benchmark to constantly try and exceed.

49:35 They they get to do their own evals on it every single day because they are the ones coding the the models themselves. >> And it's the most technical audience of all time where when they deploy an agentic system and they run into a either a bug or a problem or like some MCP server comes back with like connection invalid, they fix it. >> They know how to triage the problem. They don't call it. They just like, "Oh yeah, no, I didn't open up that port.

49:57 Sorry. I'll go fix it. So, that's like five things. Oh, and maybe like the six like it's just like a very very highpaying vertical that that like it's just like automatically valuable if you could get 10% or 20% productivity gain, let alone 5x productivity gain. So, so, so take those five or six things that coding has as as sort of beneficial properties to sort of automation, then compare that to, you know, you know, every other form of knowledge work. and you'd probably have like a histogram and I don't know if anybody's published this, maybe you can like the similarity to coding and like like what what are the domains that like like start to sort of sort of you know look closer like like less and less like coding as you as you kind of scale out. And it's like okay well lo and behold legal is kind of interesting because because like there's a lot of value creation to somebody sitting at a computer reviewing legal documents, writing legal documents, like processing large amounts of information.

50:51 Okay, so that's kind of blowing up and then you kind of go down the list. Now, let's take something like like a sales rep. Okay, so much farther down the list in terms of of sort of likeness. The sales rep's value creation is basically convincing an external customer to buy software or technology or a caterpillar truck, you know, from them. That is the value creation to the to the economy of the sales rep. >> And so, let's say we brought the world's best automation to them. Like first of all again they'd have to like figure out how to technically wire it up. They'd have to make sure they give it all their data all these kind of things. But no matter what the they're they're still rate limited and constrained by like did the customer respond to them? Do they want to meet? Can they meet next Tuesday or can they meet today? Like does the customer have budget? All of these other things. So that that's maybe the entire continuum right there is like one is like I mean of knowledge work like obviously this is not even you know touching the the sort of working with with Adam. So that's a continuum which is on one end you have somebody rate limited by so many external factors on the other end you have somebody who could sit at a computer all day long and just type type text and that is your ability to basically automate things is is that continuum. So for the real world, we have to basically bring intelligence to these workflows in ways that are that sort of somewhat feel like the shape of their workflow, somewhat feel like the shape of their work and then find a way to deliver the change management, deliver the implementation, get data into a format and into an environment that actually works with these systems. like asterisks like one of the other big things is like if you go to most engineers in 2026 like maybe like minus two months ago given the given the latest phenomenon but like like the code's in GitHub you just like connected to GitHub like like remember there's this period where like your when you launched a coding agent like there was no like sign up or register it was just like give us your GitHub that doesn't exist in knowledge work there's no like give us your GitHub for knowledge work there's like >> give us your box >> well box customers have a have a much easier time with all this unfortunately like We're only 1.3 billion in revenue run rate. So like like that means there's a lot of people not using Box.

52:52 And so what are they using? Their data is in on premises systems, legacy fileshares, legacy infrastructure, you know, enterprise environments that don't talk to agents particularly well. So just think about that sort of distinction between, you know, implementing coding agents versus everything else. even even something again you'll sort of fall asleep about is like access controls in the enterprise are totally different. And I actually totally forgot that that point about coding. In coding, you get access to basically most of the stuff ever relevant for your job. In knowledge work, you're like you're like, "Hey, Sally, can you open up that that sort of file share for me? Can you open up that that, you know, sort of project because I didn't get access to it. How do you make sure the agent has access to those set of things?" All of that work has to get done. So, the thing I think we have to prepare for is two things. one, Silicon Valley has to prepare for diffusion taking a lot longer than they think or then I'll say we but I really think they because because I like I know how long it'll take. And then the second thing is is the good news is this is sort of all correlated to the applied layer. Yeah.

53:55 >> Value creation like because the the companies that will just have the patience the the sort of the the full sort of domain expertise the the just the sheer work ethic because it's not like everybody's just from the floodgates like you have to go and and just you know pound pavement and get out there. That will be the applied layer. So I think this actually represents, you know, a trillion dollars of applied layer AI value is is actually how do you go get the technology to the lawyer or to the sales rep or to the life sciences researcher or to the person that runs the customer support team. That's all sort of opportunity right now that exists.

54:32 >> Awesome. I'm going to close by asking some advice for other founders. >> maybe let's start with founder advice and then company building advice. On the founder side, seems like you are in every AI cap table. you know, every cool new company like you know, you know, Engram, you know, how did you kind of get yourself in the middle of the AI conversation? >> There's probably two two parts. One, I was just very well primed for it like working with unstructured data for 20 years. You just like can instantly see the the the benefit of agents on that.

55:00 So like it honestly like took longer than I would have wanted that that we we got to have this conversation because we we tried to have this conversation, you know, eight, nine, 10 years ago. Y >> and now it's finally happening. So, so first of all, just super well primed. Our product sort of, you know, shape and and what people do with our product already lends itself extremely well to agents. So, like obviously we had to bet the company on that and and then lo and behold, we had the positive feedback loop of customers actually saying, "Yeah, that would actually be very powerful if I could go read every document and answer any question." So, that that's the first and certainly by, you know, biggest by kind of a factor of 10. And then the other is just like I'm extremely fascinated by the technology.

55:35 and it's just fun. It's like >> where do you learn about it? your podcast. Doresh's podcast the Twitter I mean probably unfortunately for brain cells like it's probably 95% Twitter. I have a routine where like at the end of each night I just go through the feed and it's just like I do like I I look like some sad meme probably of just like I'm just scrolling and scrolling and scrolling and like attempting to like triangulate all the information. That's amazing because it is it is the global town square for AI.

56:07 >> It is. And I want to send an emergency alert to like everybody who's like a sophomore or junior in college and just be like follow these 20 accounts on Twitter and and also join Twitter because like this will just help your career. You're either like a year ahead or a year behind simply based on your feed. And my feed is like so so wired in. >> This is the number one advice I give to people when they're asking how to get current on AI. It's like follow follow these hundred accounts.

56:31 >> It's it like but I mean virtually first step join Twitter. Yeah. >> Like you I'll still talk to to 20-year-olds that are like, "Yeah, like you know, I I I see some articles." I'm like, "What do you mean? How do you see articles? Like I don't even know what that means." Like, "Do you just like get lucky that somebody emailed you an article? Like just join Twitter. Like what are you talking about?" So, so so I would I you know, you just have to be wired in. I I enjoy it. It's a lot of fun. I play with everything. And >> what's what's your favorite new AI product? Please say Instinct.

56:59 >> so Okay. full full disclaimer, I have not done the Instinct invite code yet. simply because I have a backlog of like three other personal assistant products and >> you gota try instinct. It's >> I know I know everybody's ra I I very excited. I'm on >> and I'm not an investor. So >> Oh, you're not? Okay. So this is like totally genuine. Okay. So I I 100% will have I don't know when this is going to run, but I'm sure I will have played with it by the time it runs. I I'm in pre-release in a couple right now that I'm spending some time with. I think the personal assistant stuff is super exciting. At least in some of my use cases, it's still showing some of the limits of like browser use as an example. Like we still have some work to do there. Like probably the entire internet needs just like a CLI for their product. so we still have like there's still some blocking and tackling at the infrastructure level for these things to be totally awesome. and then and then I kind of look like your average probably, you know, AI pill knowledge worker on, you know, every day I'm asking one of five different AI systems, you know, 20 to 30 questions on just like doing research, looking for talent, looking for what is a competitor doing, what's happening in this market, how do you expand there? so so kind of that.

58:07 >> What about on the company building side? How do you, you know, 20-year-old company at this point? What's your advice for other people that are trying to reinvent their companies to make sure they have, you know, max adoption of AI, not just at the individual level, but also at the kind of company level? How do you make your business legible for AI? >> Yeah, again, some of this we have as a as a byproduct of how we've always thought about information systems in the company. So like like to no exaggeration if you have a question that you would like to ask about the the the business that has ever been documented in a form of unstructured data. So a meeting note, a project plan, a road map, a presentation, a financial document, a a financial planning session, it's 100% inbox. So like we benefit from a data architecture that already is insanely sort of tuned for because this is how we've run the company. we didn't let anybody use anything else and and so the data is very easy for us to be able to work with at scale and then of course we have Salesforce and and all the other kind of core canonical systems we've been able to kind of have I think pretty good data hygiene so agents running on top of that makes it a little bit easier to be kind of AI AI first in and how we operate maybe a couple best practices or or things that we've seen you know first of all trying to figure out where the highest leverage impact workflows are going to be and and you know trying to kind of target those so you know our CIO's very AI pill. We have a little bit of like a center of excellence on AI.

59:31 We've hired some internal AI FTEES to kind of help with these processes. >> Do you have leaderboards? >> We don't token max. we do actually have like I mean like I guess like literally we have a list of people buy a number of tokens but but it's usually to like go and inspect like okay do we do we think that's useful or is there a learning there that we should take back to some other function. so like probably our top, you know, in the top three all AI users at Box, like probably one of them is wasting, you know, half the tokens and then two of them are like, oh whatever they're doing, we need to go like do like an internal training session for everybody else. So, you know, we had this thing like two weeks ago where I was like, can you just get everybody in a room, you know, on this particular team and just like like show them how this one person is is using AI? And then like, you know, six hours later, they were like in a room and the guy was doing like a full demo of what he was doing. Shout out to Mick.

60:23 and and so like that's the that's the kind of stuff that we're trying to do, which is just like how do we show everybody what what it looks like to work in this way. But you know, for as fast as we're moving and and we are shipping at some parts of the stack two or three times more sort of, you know, actual customerf facing products. So like I don't care about how much code but like did we actually deliver more functionality the customers asking for some parts of the stack we're doing that then you'll go talk to a friend in anthropic and they're just like and you're like oh my god we still have we still have ways to go like like I you know the particular meta kind of constantly changes which is like you know two years ago you'd be like you're just using a plugin in your IDE and you're like man that's we're not going to be that like we're we're not ready to fully you know be AI first and then you're like okay everybody roll out roll out cursor and then you everybody rolls out cursor And then you're like, and then that finally happens and then you're like, you know, you're walking around, you're like, you're not only working from Slack, just at mentioning bots doing your code. Like, what are you doing? Like, like it's like we're constantly just changing what are the what are the workflow paradigms on this?

61:25 >> Do you guys have like a Slack co-worker agent for lack of a better term? >> We have a a few that that have that shape. I'm pretty excited about like like cloud tag as like a form factor. You know, you still have to again kind of get the the team construct right and the data right. but we have a variety of ways that people, you know, kind of work with agents in Slack. I don't know if it's as as sort of slackpilled as Benny Off would like us to be, or as like anthropic or opening eyes, but like we're we're heading in that direction.

61:51 >> Yeah, I think that history books will be written about, you know, the art of business in this time. Like I would love to read the new The Art of War with like everything that is happening here because I think it's pretty extraordinary stuff we're seeing. What do you think it takes to win in AI versus preAI? And like what does it feel like to be a founder right now versus when you started books? I am both jealous of and also not jealous of like like the young founders you meet because you're just like like oh to be young again and and the whole world is your oyster and you can go in you know any direction and the leverage you have you'll meet with them and the and you're like oh my god like I I saw a product a week and a half ago and they they did a demo of the product and I was like I was like absolutely this would have been a 40 person project five years ago like or or they you know especially when we were starting out easily 40 person and it was it was two people and you're just like how do you have so many tabs that work and they all seem to have stuff behind the tabs that all seem very functional like this is not fake.

62:48 >> and and so I'm very jealous of that. Like it's like incredible because you could just like you can start your company from scratch with that as the design principle. Now we will we will get there because we're just going to muscle through it. There's a couple things that we just can't do which is like we're not like we're very uncomfortable at the idea of like sort of you know removing the code review and like some of these things that get talked about because our customers can't possibly entrust us with their data security and compliance if we don't take that seriously. So we're always going to have a little bit of a discount on the productivity because of of where we are in the stack what we what we do as a business. But so jealous of of being able to kind of be fresh in that. And then on the other hand like it's like man like at the same time for every great idea it's like instantly five competitors. We didn't have that, you know, problem. We had a we had a good kind of couple years where we just could like we could just grind on our product and our and our sort of experience. And it wasn't like every 3 days you were like, "Oh, like Sequoia funded this thing and benchmark funded this thing and like you know, we weren't going kind of nuts like with that." Now we we had our own version of that. So like at the time I probably was going going nuts, but it was like like in retrospect it was like not worthy of going nuts. Now it's like, oh man, this is like a real race in every one of these markets.

63:58 >> Totally. >> So I think I'm probably pretty consensus on this, which is in a world where AI builds things so much faster. Then probably the shift moves to whoever can actually get it to the customer is in the best position. So you know, it's it's fun talking to founders that are pretty kind of pled on that. you know, Scott or Matan are like they they get the mandate. They're just like this thing is going to be an enterprise diffusion play. and and so you have to get it to the enterprise. and and anybody who kind of mistakes the the mandate right now, you're just going to lose. It's just like it's game over.

64:32 Sorry. Like there is quite literally trillion to trillions up for grab at the applied layer and the companies that that know how to build the teams and and get to the enterprise will be the ones that win. It's just like obviously guaranteed. >> Well said, Aaron. This is a very fun conversation. Thank you so much for joining. >> Thanks for having me. >>

Summary

Aaron Levy, founder and CEO of Box, discusses the transformative potential of AI in enterprise workflows and the importance of being "wired in" to relevant information sources for career advancement. He emphasizes the need for companies to adapt their products to leverage AI effectively, bridging the gap between AI models and real-world applications.

- Levy advocates for college students to engage with Twitter to stay informed and connected in their fields.
- Box is evolving its platform to incorporate AI, focusing on automating workflows and enhancing data management.
- The application layer of AI is gaining traction as companies seek to bridge the gap between AI models and enterprise workflows.
- Levy highlights the importance of understanding specific customer needs and workflows to create effective AI solutions.
- The conversation touches on the challenges of integrating AI into existing enterprise systems, particularly regarding access controls and data management.
- Levy notes that the speed of AI diffusion in knowledge work is slower than in coding due to various external factors and the complexity of workflows.
- He emphasizes the significance of building a strong internal culture around AI adoption and leveraging existing data architecture for better outcomes.
- The future of enterprise AI will likely involve more background automation, with users receiving results from agents rather than initiating every task themselves.

Questions Answered

Why should college students engage with Twitter for their careers?

Engaging with Twitter and following key accounts can significantly enhance a college student's career prospects by keeping them informed and connected.

What challenges do enterprises face when moving to the cloud?

Enterprises need effective solutions for securely storing, sharing, and managing unstructured data as they transition from on-premises systems to the cloud.

How does Box evaluate the effectiveness of AI models?

Box conducts extensive evaluations of AI models using domain-specific tests and real-world applications to ensure accuracy and effectiveness.

What is the significance of systems of record in the age of AI agents?

Systems of record must integrate effectively with AI agents to enhance productivity, rather than creating standalone solutions that complicate workflows.

What are the key challenges in automating knowledge work?

Automating knowledge work involves addressing unique workflow shapes, data formats, and change management processes that differ from coding environments.

© transcribe · For agents Built with care and craft by Gokul Rajaram