Transcript
0:03 We got to change the game. Earlier this week, we probably had three customer meetings uh that were all really, really good. uh some key customers app say and this weekend this weekend uh I went on something called an SF Hills ride which I've not done in ages. It's a 5-hour ride through all the hills of San Francisco. Um >> it's a bike >> it's a bike ride through all the hills of San Francisco and that was enormously fun. I was not sure if I was fit enough to do that but ended up being a lot of fun.
0:45 >> So that's that's my last few days. >> Oh that's that's amazing. I also saw your um the webinar you did on Tuesday. >> That's right. The panel discussion or the webinar, sorry. Um was was that yesterday, I think. Yes. Or uh I'm confused. No, day before yesterday. You're right. >> Um >> no, that was really good fun. Enjoyed it thoroughly. >> Yeah, that was a great conversation. [snorts] So, I talked to Joe obviously, so I've already gotten the clockwork pitch and then I was able to listen to it and it really a lot of things [clears throat] started clicking for me. Um, >> nice.
1:21 >> And for me, like I've been selling infrastructure my entire career, >> right? >> So, this is cool. >> Exactly. I mean, you know, there's anything really cool in a while, you know. >> I completely agree, Sesh. I think, you know, it's the last 15 years. Um, we've become so used to the cloud abstracting away all infrastructure and almost basically not having to think about networking, not having to think about storage, not having to think about compute. It all comes from so you're really operating above that layer for the most part. And then the world of AI and AI infrastructure, data centers are back with a bang. I mean hard. It's actually one thing I will tell you when we talk to um so many of these neo clouds or even people trying to put up onrem infrastructure the skill sets have disappeared uh people that really understand these domains don't exist anymore. I mean I mean of course they do but not in anywhere near the numbers needed to fund the kind of infrastructure buildout we're seeing.
2:23 >> That was one of my big takeaways. I I have to go back to building data centers. >> I know Chris. [laughter] >> Yeah. Chris used to be a CTO uh building data centers. He's built it what 30 300 ft uh to tell them about the one that you built. >> I I built an EMP shielded data center 250 ft beneath a mountain. >> Oh my gosh. I'm curious what was in there and why it needed to be EMP shielded. But that sounds cool. So, it was some government work and I imagine uh it was an insurance company actually that wanted because they wanted to, >> you know, like we're like, "Well, there's not going to be much for you after the EMP to do anything." They're like, "Yeah, at least we'll have the, you know, we'll have this the uh raise with all the uh claims data."
3:08 >> Okay, you you do you. But um but it was wild, man, cuz I was it was you know this facility was absolutely massive and it it wasn't the the facility itself was a full acre underneath underground. It had its own zip code. It had its own fire department. Um it had a lake in it that was the chill water plant for the whole thing >> and the lake kind of ran off the back side of the mountain. So there's no never any risk of flooding with it either.
3:35 >> Right. Right. And the mining engineer who who managed the whole thing, he would he would get into this boat with, you know, big hip waiters and stuff like that and get in this boat with a a flash flashlight and a can of spray paint and he'd be like, "If I'm not back tomorrow, come look for me." >> My gosh. >> They [laughter] know where the lake everywhere the lake went. You should you skills will be in high demand, Chris, at the moment.
3:58 Well, and it's funny cuz like the a lot of the guys who built the um the you know the the black site in Utah for the NSA. >> Yeah. >> Same same crew who was built doing the EMP shielding >> and they man they they they got they made me think twice about a career in security. >> Like I don't want to I don't want to live with that kind of paradigm. >> Yeah, I know. I know. Exactly. Exactly.
4:20 I'm sorry. >> Yeah. It's funny like how how it all kind of comes back around, right? like you said from the infrastructure folks now going to the cloud. So now those skill sets are not there but there's also this uniqueness we won't get into it >> the uniqueness of high performance computing these these these GPU powered environments to in my mind and I think probably most people they just think what it's hardware it's servers and it's CPU is you know chips and it just works and there's storage and there's there's been networking around for years what's the problem right um and so like what's cool for me is like there's new problems to solve, you know, >> for sure. I mean, oh my gosh, first how much the energy density of the data centers that you're putting these into is totally different. So, where do you get the power from? Then, how do you the racks are basically dissipating a lot more um heat and they have a lot more power draw. So, you have to change the racks all over. The cooling has to be completely redone. The cabling per server is roughly uh 8 to 10x the connection density just because I mean where a traditional server might have a couple of nicks, these modern servers have 10 nicks in an 8GPU server and far more in a 72GPU server. Um and so and these are roughly operating at somewhere between four to 8x the bandwidth of a per on a traditional server, right? So, so you start doing the math on every front. Uh that's part of the reason why failures are so rampant uh in in this environment because almost everything is an order of magnitude faster, harder, denser and so that means just things are just operating at the edges all the time.
6:14 >> Yeah. Yeah. Yeah. Well, I mean I I I remember building, you know, like we would home run fiber and and you know, like all the like these big single mode bundles would be like this big around and they like make the the racks come down from the ceiling cuz they were so heavy. Exactly. The problem is you can never ever remove a cable because >> where's that cable go? I don't know. >> That's right. No, that's exactly right.
6:35 That's exactly >> No matter how well you label them, you still just can't trust it. >> Yeah. >> And pulling them through is going to, you know, can disconnect things and break things too. Exactly. We have one large customer where our contract we typically price on a per GPU basis and per agent or per node basis and they basically priced it on the base of uh where we are deployed what's the kilowatt uh of the racks in which we are basically deployed if you will and so it's almost like the unit of scale has become how many kilowatts or megawatts are uh are are basically able to be deployed in in a in a data center context and and so It's completely u the constraints are quite significant on a lot of fronts.
7:18 >> Yeah. And the cooling is the hardest part. >> Yeah, for sure. Like [clears throat] indeed you get the power there. It's scary. I mean like some of the power they're putting into these small >> water cooled systems and crazy cooling structures. Exactly. Exactly. >> Exactly. And then you start some peeling the onion on, you know, I still [snorts] remember the days when pabytes were massive and now you're just starting to see more and more of these large scale data centers talking in exabytes as if it's just the I mean it's sort of in a heartbeat. We've gone from terabytes to pabytes to exabytes. Um [snorts] and so the amount of data, the network bandwidth and the speed latency that you need just every every angle of the infrastructure design, you're pushing the boundaries. That's what I that's what's make making it fun because in some sense oh gosh um >> I don't know if you recall right when before virtualization the idea that servers were operating at 30 40% utilization because you could really virtualize the application it was a single application per server um was the mindset and then virtualization came around and you could cram six or seven VMs onto a single physical machine you started increasing the utilization we're back into utilization for other reasons, not for very different reasons, but we're talking about utilization in the 30 40% zone. We're talking about availability that in the cloud is measured in the for compute 49s or 59s.
8:43 We're not talking about 92 93% uh availability for for the computer instance because most the rest of the time it's just there are failures that have taken the system down. On some dimensions we are pushing the boundaries so hard on other dimensions we feel like we have to solve problems that have been solved all over on things like resilience, fault tolerance and and that is how that is. So like I mean at some point we we used to talk about like tier three and tier four data centers. Right.
9:10 Right. >> And like to have a tier four data center you have to have staff on hand all the time that are HVC AC engineers. You have to have electrical engineers. you have to have like that those skills present and available all the time you know in addition to all the you know two end pathing and stuff like that and and you know like we were seeing the change and it was like there's no it doesn't make any sense to do that it makes more sense to just make your infrastructure >> more fault tolerant and that that was that rise of the S world where like and then it's like let's just design to what the capabilities of this platform are >> and and build you know to accept that >> and that's where I think we we've kind of landed And so like those failures are sort of built into the system.
9:52 >> Exactly. And that's what that will happen here too. We're just in such an early innings on rethinking infrastructure for AI that the first focus has been speed right now. We're starting to say well security matters, resilience matters, observability matters, um automation matters. And so because in the end those things actually don't take away from speed. they just make it possible for you to operate in a much more predictive predictable manner. And so I think you're starting to and unfortunately part of it is also the technologies are needed to achieve those very same things whether it's fault tolerance or deep observability while the while the discipline is the same the technologies are very different. Um I mean the the distributed communication protocols in training are not built naturally with the ability to do sort of a retry after a timeout automatically. Don't let a failure in the network propagate up to the application layer. Those are well understood principles of protocol design but that's not necessarily how training pro distributed uh training works uh today. Right. And so so there's many things that I think will evolve for sure.
11:04 >> Chris, I you know, you should talk to um I I talked to uh Vjoy Pandi at um he he runs uh he's like the VP and general manager of Cisco Outshift, >> right? >> And I don't know if you're familiar with Cisco out, but it's sort of their like internal incubator. >> [snorts] >> And so like I'm not actually the two projects that he's working on like sort of longerterm horizon projects. One is one is sort of quantum networking which is really an interesting thing.
11:34 [clears throat] They've got a they've got a chip now at room temperature that'll entangle 200 million photon pairs per second. >> Wow. And they can send it over like normal fiber optic and then do like quantum teleportation and things like that because the problem is like with with quantum computers is like you can get to a thousand cubits but you need like 100,000 or a million and it's hard to do that. So like they're say they're voting they're like gambling on like well let's put you know 100 [clears throat] thousand cubic computers together with quantum network and they just make it easier. But the other the other pro problem they're they're focused on in in the shorter horizon now >> is uh uh >> a network of agents. That's what they call it.
12:18 >> I see. >> So they're they're building all the new protocols for agent inter agent communication in an AI world. So they're doing all that work to tie it all together like kind of some of the like the protocol level stuff. They're like kind of rewriting the whole stack >> to like integrate AI and inter agent communications. So you should re, you know, you should reach out to them. I can give you >> No, I'd love to. Yeah, it seems like both. I mean, gosh, they are sort of bleeding edge uh problems to attack.
12:44 >> They're putting together a a working group around this. >> Okay. They're very helpful. Very helpful, especially given the focus on communication. >> Yeah. Yeah. I mean, I I'd love to put you in touch with that with them. >> I think we're going to change the name of this podcast to Nerd Talk. Nerd talk. I think >> I love it. I love it. [laughter] >> I love it. I knew you two would. I knew you two would.
13:09 >> Well, you got you got to love a nerdy CEO. >> Yeah. >> You [laughter] don't you don't encounter that too often. >> And that's that's the fun of being at a startup. Otherwise, if uh otherwise you might as well join a large company. It's a it's that's um at least that's what I thrive on. >> Yeah. No, >> so I'm just going to get started here. We're already recording and the reason why is because of what just happened. I know that we're going to get some great clips and uh our team will edit it and we we'll give you everything. Um >> people will be looking up Cubix and whatever else you guys are Yeah. Oh, I see.
13:44 >> Although now you just named he created a new company. >> Exactly. Right. There you go. Exactly. [laughter] >> I am the snails in marketing guy. >> You squint and there'll be a new technology based on cubics instead of cubits. I'm sure >> Cubix's the quantum quantum AI company. >> Yeah, [snorts] exactly. >> Smart people will be like, I wonder what that is. >> Yeah, >> yeah. So, uh I want to ease into this. It's going to be very very casual. Um but I'm going to I like in the beginning to give you some softballs.
14:13 >> PE what we're learning is people want to know what you think and you know part this this podcast we're calling it what the future because that's what we're trying to do is figure out what is the future. Right. Right. So, um, just couple questions around the AI bubble. I want to ask you about that and then just a little bit about what Chris was just talking about too is like you're a technical CEO, but >> uh, you know, I will I will kiss your ass here in a second and brag about how awesome you are, but like you really >> I really respect you as a leader mostly because of what people say about you. To me, it it's off the charts. So, there's a lot of speculation these days on if we're in an AI bubble. What's your thoughts?
14:58 >> You know, I think I'll say two or three things. Um, one, I think we have to separate out what's happening in the stock market from the technology side of it. That's the first thing I'll say because I think on the >> technology side of it, >> I think it is one of the most >> incredible developments in uh the last thousand years in terms of transformative power. I do think people that say the AI revolution is like the industrial revolution are not at all overstating the case. I think what we are on to is the beginning of something that will change everything about sort of how we live, work, uh, and do day-to-day stuff. Uh, some good, uh, some bad, but it will be transformative for sure. That's the first thing.
15:43 >> On the stock side, um, you know, almost all great big uh, changes have have been accompanied by some excess followed by sort of a reversion to the mean and then some sort of a couple of hype cycles if not more. Um, and so it's possible that we are going we're going to have a stock market sort of crash for AI valuations in in particular. That said, here's one thing that struck me. If you look at the amount of spending, let's take Nvidia for a moment and say if you look at Nvidia's revenue and say how much of that revenue is coming from companies that are using other people's investors dollars to buy Nvidia product, which would be the classic sign of a dangerous bubble. Um I would say a vast majority probably 75 80% of their revenues are coming from really large cloud companies that are able to fund their annual capex out of their annual cash uh cash generation. That is sort of so that almost puts a floor on how much frothy revenue exists here that's made that's not coming from people with real operating cash flow. So it's possible that the stock market valuations will come down. In fact, they look super crazy [snorts] uh high. That said, I I also think there's a floor here because I think there are real business models underpinning the purchases of AI infrastructure.
17:06 >> Yeah, I was going to say I think it's really interesting to see like what happened with Amazon recently where, you know, they had this massive layoff, but the and and they they sort of pinned it on AI, but when you dig into it more, it looks like they actually laid off all those people so they had the money to buy more GPUs. Exactly. I mean yes. So when I said so not all good I do think we'll go through a um sort of a period of turmoil when it comes to employment and sort of reskilling and there'll be some so jobs displacement is one of the penalties of what will happen over the next few years. And so on the one hand you're seeing companies make completely uh uh unprecedented profits compared to any prior year. And those very same companies are also saying we have too many people um in our payroll and we need to shrink the payroll. And so so that that is unfortunately uh both happening at the same time.
18:01 >> I think it's very funny like we talk about people being replaced by AI but I mean these people are literally replaced by a GPU which almost [laughter] seems >> exactly right. It's actually a little bit more insulting. >> That's right. [laughter] That's right. That's right. >> Oh that's classic. >> That's right. That's right. Uh well on the subject of people so you are seen as a great people leader and I remember those days at Netat when I first met you there and you were just promoted promoted promoted you became CEO of Nimble CEO of Cy. You you've had an amazing career. Um, but what I really appreciate is as every time I've met you and been with you, it's always been such a pleasure, easy conversation and um, like it feels like you care, you know?
18:47 Um, and I think that was a lot of the the NetApp kind of culture, right? >> No one cares what you know until they know that you care. >> So, I'm curious what how would you describe your leadership style? uh balancing people but also the rigor and to be tenacious in this market. >> Well, and I think you got the passion too which is a really big part of it as well. >> No, I appreciate that Chris. I I think I'll start by saying something sesh. I honestly I feel like uh NetApp was basically a 10-year school I went to. I learned so much at NetApp that even sort of dimensions on on how to lead companies or organizations or teams. I feel like I learned a lot from mentors like Dan and Tom at uh at NetApp. And so NetApp I credit a lot for sort of uh over the years how um I've at least been uh leading organizations and teams. And I think there are a few simple principles. I I do think the starting point um for me is is is almost a belief that the team actually can accomplish a lot more than any talented individual.
19:53 And so to harness the power of the entire team, you have to essentially set great context. So, so real transparency in communication, make sure that you're so I typically do 50 to 60 one-on- ones uh in my last company, Sydney, where we had an employee base of 600 every month. Sort of 50 to 601 ones said. try and talk to as many people establish context uh and have a uh very free sharing of information so that there's great transparency in the entire culture.
20:22 That's the first thing that I would say is extremely important and and ultimately the goal is a deep belief that [clears throat] a group of individuals that may be uh individually not as smart as one particular talented individual will always do more if they can work sort of uh with really good shared context. That's the first thing I would say. The second thing I think is a clear sense of purpose. I feel like organizations need to know they're working towards something and and frankly in tech companies that sense of where you're heading has to come from a technically grounded perspective of what is the vision that we're working towards. What is the when we succeed what will the future look like? So let's paint a picture of the future. And in a tech company, that picture of the future has to stem from almost an understanding of where the technology will go, how we can intercept it, and then getting people excited around it. And so that's the second thing I would say is just imbuing the entire organization with here's what our destination looks like and why that's a good place to go is is probably the uh second thing that I would say is key. And then the I would say one other thing is, you know, um culture really matters, right? Right.
21:34 And so what so what what I mean by that is um these are about where you go. Sharing of information is how you operate but but also what is valued in the company. Uh at SISD we had a princip sorry at Nimble we had a principle which was basically no jerks right and so it doesn't matter how smart you are. If you're not fun to work with then you're probably not going to make for a great team member. And so the culture matters and it's things around how people are rewarded. what is sort of lorded in the company versus what's frowned upon in the company and so essentially what you encourage as behaviors within your team within the larger organization and is that creating a collaborative environment so that's probably the third thing I would say and that sesh you know this I I absolutely saw that very much at NetApp and and that's something that stuck with me >> you know when you you mentioned sort of the vision part because I I totally agree with you on all these things and I think what I see leader ers struggle the most with I think is communicating an effective vision >> and and sticking to it. You know, because in a lot of organizations the lack of vision because if you have the vision, everybody marches towards it and that's you don't have to tell people to do stuff. They just kind of know where they're going. That's right. That's the best way to do it. That's right. Like what what what are the tricks to like really communicating that? I mean, I guess one of them just have a good vision to start with, but >> Yeah.
22:57 >> Yeah. So I think the uh uh there's there's a couple of things that I would say because I think um having uh a vision that is um the clearer it is in your head, the more I think you should be able to translate that to multiple layers of the organization. the same vision when you're talking to the finance team and how you're communicating it or when you're talking to a business decision maker uh on the customer side and how you're communicating it is very different from sitting down the in a room full of engineers and architects and how you're communicating it and so I think and and I believe the ability to translate up and down certainly comes from having good communication skills but it also comes from clarity around the vision itself so I think that's the first thing I would say is when you have a vision find ways of communicating it even if it takes explicit work to do so at various levels. A one-pager all the way down to a deeply technical document that describes the architecture of where you want to go has to be very very uh that has to exist in the organization.
24:03 >> Yeah. >> The second thing I would say is Oh, sorry. Go ahead. Sorry. Go ahead. >> No, no, no. Please go on. >> No, no, no, no, no. Second, second thing, please. >> Yeah. Yeah. The second thing I would say is is uh you know so much of our energy is spent on how do we communicate that externally whether it's in our first call decks whether that's in our website messaging etc. I think more of the battle is is really in synthesizing it internally to employees and making sure they are bought it. So find forums right whether it's written forums all hands meetings uh team meetings etc. And actually this is something I I did not do uh really well and I and I continue to sort of struggle at this is you feel like sort of if you've communicated it five times you've done enough or if you've done it 10 times you've done enough. And I find gosh there's no amount of repetition. Find different flavors of saying it. Don't say it the same way each time but but you have to keep coming back at it every new product. Let's go back to our vision and why this new product fits. every new customer that we want, let's go back and talk about why did this customer become a customer because they bought into our vision and here's how it fits into it.
25:09 So, I feel like that's something that you have to do also consistently articulate what your vision is in a in a way that sort of you don't feel like you've done it if you've done it a few times. >> Yeah. >> Yeah. You're Yeah. So spot on. Um the other the third point that you made about culture, you know, for Chris and I that like hits right in our heart, you know, because we're indeed >> um it is uh well, we were spoiled at NetApp, right? I >> I could not agree more. I you know, it's a I I joined NetApp from a u from another company u Mckenzie, the consulting firm. And just the honest truth is there's there's been so many scandals around McKenzie lately that sometimes it's hard to remember. But when I joined McKenzie, uh it was it was I think there was a book comparing it to the Jesuit uh priests or something because they were so sort of evangelical about what their mission was and so on.
25:59 I was very proud of the organization and I I still am not withstanding sort of the noise around Mckenzie. But part of what I was going to say is another place where the entire product was people and the expertise of the firm, right? There was no technology they were building and so they had to uh it's a different style of culture. NetApp's culture was beautiful in its participative nature etc. Mckenzie's culture was all about sort of um fostering talent at the expense of everything else, right? So sort of how do you take this brightest minds and motivate them and sort of keep making sure that the organization maintains a high bar on talent. But what is common to both of these is ultimately how do you take whatever it is that you believe in and make that percolate through the entire organization.
26:42 Whatever your culture is, how do you reinforce that and make sure every person's living that? It's almost a that's what NetApp did so well is sort of they it was a strong culture but how did every employee whether it was the hundth employee or the thousandth employee seem to embody that culture I think they did that really well >> and we the other thing that helps culture is when you have a very successful sales organization from >> sales team architects and anyone supporting >> the sales team >> if they feel they're winning if they're coming in to work every day they believe in the vision They're excited about people.
27:28 >> They I firmly believe a lot of people don't stay at these companies because of money. >> They stay there because of other factors. >> Indeed. Indeed. >> I know for me that's definitely the case is the people. >> No, for sure. practices like Tom calling not just sales people who close deals but engineers that helped close deals and support people that helped support the uh customers that ultimately did expansions those allow everyone to focus on success rewards everybody right and so I I agree >> yeah I got to tell you this quick story and then we'll get started here on the Tom Mendoza topic so uh a really good family friend of ours my wife's best friend's daughter they live in our neighborhood here >> and um she got into Notre Dame and her >> You know what? Before you go there, you should explain who Tom Mendoza is.
28:14 >> Well, I think >> not everybody who's listening is going to know who Tom Mendoza is. >> I think our our audience definitely does, but Tom Mendoza was when I started in '98 and NetApp, he was the director of sales and marketing. He was then became president and vice chairman and now he's on a bunch of boards and he is just he's all >> Mendoza School of Business. the Mendoza School of Business and just that story when he tells you about when how that all happened. It's just so awesome >> indeed. And Tom and Dan were were sort of great partners in shaping that entire company's evolution.
28:50 >> Yeah. Yeah. Totally. >> For sure. >> So, so, uh, our our family friend's daughter is going to Notre Dame. It's freshman year. She's a little, you know, just like anybody would be a little nervous. And I just reached out to Tom and I just said on LinkedIn, hey, you know, she's going to Notre Dame, you know, as a freshman. She's super excited. She, you know, wonderful person, wonderful family. Uh, dad went to the Mendoza School of Business, too.
29:18 Um, would you mind just sending her a note? He's like, "Send me both of their phone numbers." And, uh, he created a voice uh, sorry, a video for each of them >> and sent it to me. >> Oh my gosh. So it's just amazing that way that is he makes time for these things. That's what is amazing. >> So important. So important especially these days in the world of work from home and >> you know social distancing and distractions of these phones and devices. You know I'm I'm very very very big on the humanity side. That's why don't >> enterprise AI tech sales and humanity cuz >> you know we can't forget that's why you know we're all humans.
30:01 >> Agreed. >> Agreed. Agreed. Agreed. >> Agreed. >> Well, let's dive in a little bit. So, Sur um >> for people that might not know you, can you just tell us a little bit of history of, you know, your from, you know, your young ages to kind of brought you Yeah. What brought you here? Uh how' you how did you get to where you are now? >> Yeah. Absolutely. So I I was born and raised in India. Uh did my engineering undergrad and u uh business school in India and then I joined uh McKenzie the consulting firm in the early '90s and towards the late '9s I transferred to the US to Chicago in particular with McKenzie >> with the intention of going back after about a year of a short-term stint. But that's when a friend of mine had joined NetApp. uh and having when I visited him, it became clear to me that the excitement of the late '9s, this was 98, the excitement of the valley was just completely uh impossible to ignore and get drawn to. And so I left Mckenzie to join NetApp and I almost spent 10 years at NetApp. That was really in many ways the beginning of my tech career and tech education. and and and I joined as a an individual contributor managing one of our alliances with Dell at the time. Uh moved gradually into product management and then onto the exec team uh to run engineering and product management. And so that was a one of my most that fun 10 years uh uh stints where I learned an enormous amount. The last two years were interesting in that while I was running uh R&D and product management, I had just come off of a year and a half of running a business unit that we had acquired, a company we had acquired.
31:45 That was my first virtual CEO experience because this is a security company called Dick Cru that NetApp had. We decided to structure it as an independent business because it was partnered with EMC and other storage companies. And to give it a chance to survive, we needed to think of it as an independent sub with its own sales, own marketing. That never panned out. When EMC acquired RSA, the idea of an independent security subsidiary that would partner with EMC did not make any more sense. And so we merged it back.
32:15 But the one and a half years of running a full business sort of unfortunately spoiled me in the sense that I did not want to go back into a functional role after that. I loved the idea of running an entire organization and uh the business uh not just the product. And so that's when I left and since then I've just been doing startups. Um Chris and Sesh I was at um my first startup was a company called Omnion. Also an interesting experience this when I joined they had an S1 on file. dumbest uh timing wise, dumbest uh time to have left NetApp to join because it was uh 2008 uh towards the end. So when I joined the the revenue was something like 37 million uh when I signed the offer 8 weeks later when I joined it was about 23 million uh and so really the bottom had fallen out of the market. We had to pull the S1, we had to restructure the business. Um, but ultimately uh we were acquired and it had a successful exit and then I think began what I think of as my probably my most fun um CEO startup experience maybe organizational experience. I joined uh Nimble. We were about 20 Nimble Storage.
33:23 We were about 25 people or thereabouts when I joined. Just had launched the product. I'd been on the board for a year prior to joining as CEO and so it was really pre pre-revenue and over the course of the next six seven years the company became a public company and then went on to be acquired by HPE just about a month ago I had the entire Nimble team at my home for dinner so it's a it's more than sort of the milestones that are business milestones I remember it was another place where the culture felt exactly like the NetApp culture almost everyone I meet from there talks fondly about their time at at Nimble >> for sure.
33:59 >> I then left um took a year off and uh wanted to do something completely different from hardware um and um infrastructure which is ironic considering where I came back to but went to a company called Sysdic partly enamored by the idea that Kubernetes was going to become the basis by which cloud applications were built. Um Sysdic was really an observability and security company for containerized microservices. Um again joined pretty early. So I think we were about 80 people when I joined just under 5 million in in ARR. uh I left about 6 years in we were we're now SDI is now uh about 120 million 600 people successful very uh strong in container security uh an open-source pioneer in in sort of cloudnative uh container security but for a couple of reasons um mostly to do with AI I decided it was time to leave cyics so just the last part of my journey and and arriving at clockwork for the last year or so when I was at Cydic, we had already started to embrace AI to try and bring agentic uh security to bear within the cydic platform. And I I'm I'm seeing firsthand how transformative AI can be when used well.
35:21 I've also realized as an existing company that has existing customers a pre uh existing architecture that you've built on you can do a lot to embed AI and Cydic is doing that really successfully but two things struck me one the next decade and a half every product is going to be designed ground up with AI in mind rather than embedding AI and the second thing that fascinated me was as much as you can think about companies that use AI I I was fascinated by the the underlying infrastructure and the underlying models, the underlying actually foundational layer itself that was then feeding these companies that were using AI. And so I decided it's I was all actually going to retire after sisdig or do something a little less intense uh after 17 years of startups.
36:12 But the draw of AI was so strong. I wanted to go back to doing something uh in the AI space and ideally a startup which is always um sort of my preference. And so that's that's how I came to Clockwork. And I'll talk a little bit about I knew the founders of Clockwork and so on, but that's what brought me to sort of Clockwork in many ways over >> that the lure of AI is is incredible because it it is such a it's a unique period of time. I've never seen anything like this in my all my years in it.
36:41 >> I could not agree more, Chris. It's a I was on the board of Clockwork for uh I joined the board last year and then uh about six seven months ago is when I came on full-time as CEO. But it it's so it's a it's a >> it's hard to uh exactly characterize how so much change can happen in just 6 7 months of even within clock. I'm just seeing the pace. I've never seen anything move as quickly as what we're seeing. Whether that's at the infrastructure layer, whether that's at the model capabilities, whether that's in terms of how companies are embedding AI, it's just wicked fast.
37:20 >> Yeah, it's going to be an interesting hype cycle, right? Um, [clears throat] >> for sure. >> We we've seen these hype cycles before. There's just something very unique about this one >> on one side. I don't think it's just I don't think you can look at AI as like one thing on that hype cycle because I think there's things that are in the trough of disillusionment, peak of inflated expectations and the plateau of productivity all at the same time, you know, because there's so many different things going on here.
37:44 >> Sure. No, no, I get it. At the same time, I think from the people that aren't in the AI world, um, they're in panic mode. And >> for people like us that are in it, >> Yeah. >> Yeah. I understand why they have some panic, but we'll be okay. We've been here before. This is just the part of innovation. You know, the it's it's everything's going to be AI. You know, back in the day, if you were if you didn't have a cloud story, you weren't getting funding. You know, if you weren't going SAS, that was going to be a tough road, right? And now it's if you don't if you don't have AI, you're not going to get funding. And, you know, no one really cares. U but ultimately, I do think this is the bubble is going to burst. Um, and I think a lot of startups will fail, but I think that's a good thing from the perspective of we're gonna learn so much >> in the next Well, we've already learned so much, but this is not done, right?
38:38 Like, we're going to get into this like some of the problems that are now out there that we need to solve. >> Indeed. Indeed. Indeed. It's absolutely amazing, you know, and some of them when we talk about them, I they're almost elementary to me in some ways, but they're so incredibly impactful, you know, >> for sure. I I think we are early in optimizing everything that goes into from the infrastructure up to sort of agents that are delivering services to end customers. Everything is still very very early in its uh evolution.
39:12 >> So, uh why clockwork? what brought you there and maybe if you can give us a little bit of that history. I mean I know it but you know a little history of how did clockwork actually become a company. >> Absolutely. So I specifically I uh got drawn to Clockwork partly because of the problem they were solving or or in large part because of the problem they were solving and the technology they brought to bear. I happened to know the founder Balaji who's a professor from Stanford.
39:38 I've known him socially for the last couple of decades off and on. uh and then as as as I understood a little bit uh the co-founder Balaji's co-founder Yilong uh is in many ways sort of the the thesis that he worked on with Balaji at Stanford is the foundation of clockwork and as I understood more of what was the technical foundation of the company I got completely enamored so let me talk a little bit about sort of the evolution of the company itself fundamentally the company started when Elong created this um mechanism to synchronize clocks on a uh large so within a data center across data centers across a large thousands of machines using purely software where the clocks are synchronized to within tens of nanconds of each other and for the and so when you get accurate clocks in terms of drift between machines there are the early four five years so the company was started in 2017 for the first 5 years a lot of the use cases frankly was operating as a uh Stanford group if you will. So more as as an offshoot out of Stanford rather than as a full commercial company. It was a small team of four or five people. And really the use cases were vertical market use cases. So we have companies that are Fortune 100 financial companies that want to time stamp their records so that they have accurate timestamps. There are companies that are uh crypto trading companies that are using this to make sure that they can optimize um how to place a trade in the most efficient manner from a timing perspective. So there are a group of customers that were really vertical market customers and that's where the company was focused for the first 5 years. In 22 and 23, the first big leap um that happened was applying the fact that when I can synchronize clocks accurately, when I can accurately measure one-way delay, the time it takes for a packet to go from machine A to machine B or virtual machine A to virtual machine B or container A to container B once. So I can measure that delay accurately and I can measure oneway delay. So if time from A to B is different from B to A, most people basically estimate latency by taking roundtrip time and dividing by two. And so we're able to get accurate one-way latency. And then we built an entire uh network telemetry portfolio around that. That's the first time that the company leaped from being focused on vertical market timestamping applications into how do we optimize uh networking across large number of distributed machines, if you will.
42:16 The uh step beyond just measurement came when we were able to embed control logic. So a software control plane to say this is all still focused on TCP networks connecting containers and virtual machines. What we were able to say is now that we know the delay between machines using software and sort of relying on the network itself can we do things like congestion management. So if there's a congested uh set of links then let's slow down the pace of traffic so we can manage congestion. Can we do quality of service? So for example, if application A is more important than application B, let's prioritize traffic for application A over application B. So those two things, right, which is sort of taking our clock sync foundation and we built what we call dynamic traffic control, which allows you to manage congestion and quality of service on networks, all applied to VM and container uh clusters, if you will. 24 is when the big breakthrough happened.
43:13 That's when I s sort of started engaging. Up until then, we were still on the CPU side. And what we realized was that GPU clusters used for AI training are the most demanding distributed applications that have ever existed in history. They have lots of unique properties. Um, one GPU that's slow by a few seconds will make every other GPU in a distributed training job wait for that one machine. And so there's a whole bunch of properties where timing is extremely important.
43:43 All computation happens across a large number of machines. So if we can take everything that we've done with respect to measuring the condition of the network, optimizing the flow of traffic on that network and thereby improve what happens with GPU clusters that is a big transformation. So that was really the opportunity that came to clockwork in 24 and we started working on that problem. That's also when I sort of got completely excited. I I've the the I'll talk a little bit more about what we do, but that's sort of long's explanation of how I came to Clockwork and how Clockwork itself has evolved. Um Sesh.
44:18 >> Yeah. >> Well, and the problem you're trying to solve though is is kind of a non-trivial problem within the the GPU cluster world. I mean, could you speak to like like how much utilization, how little utilization we get out of these things? >> No, absolutely. So if I if I describe the problem in terms of its business impact first, right? >> Yeah. >> Typically in these large GPU clusters, um really well-managed clusters will get up to 50% in terms of utilization. Often utilization runs in the 30 to 50% zone.
44:49 Utilization being measured as if I have a thousand GPUs in my GPU cluster capable of delivering X amount of flops, then what I'm really realizing is somewhere between 3 to 0.5 flops, right? And so that's one this the within buried in that are many things. But the other really egregious problem is that if I look at a th00and GPU cluster for every 24 hours that I'm basically operating that cluster I lose between 2 to 4 hours of availability on that cluster where something disrupts my training job I have to stop the training job and restart it from a previous checkpoint.
45:24 And so I'm losing 2 to 4 hours which of course of course feeds into your um cluster utilization but it also feeds into the fact that your jobs are taking longer that basically some number of people on your observability team have to quickly find out exactly what happened and correct it. So observability so there's an emotional aspect when failures happen that's as daunting as the lost utilization. Right? No operations team wants to feel like things are always failing. they don't have complete control. So that's really the nature of the problem. Low utilization, very high failure rates.
46:01 Every incident takes long time to detect and remediate and therefore your training jobs take much longer to complete. As much as two to two and a halfx longer than in a theoretical best, if you will. And so that in a nutshell is the problem. >> Yeah. And you Oh, sorry. >> No, please go on Chris. I was going to say like I mean there's a reason why Nvidia is a you know $5 trillion company now these things are not cheap so like utilization has a huge economic impact right >> for sure you take a $100,000 100,000 GPU u cluster you're probably spending5 billion on that cluster as a in terms of that data center >> so there's a capital that you can think of uh and that's sort of wasting two to three billion is just non that's crazy, right? Crazy in terms of um sort of how much opportunity exists to improve. But but sort of equally significantly that same data center will consume uh let's say 125 to 150 megawatt of power and that means you're basically throwing away something like 50 to 75 megawatts of power. That's 25,000 homes that that basically can be lit up all year long, right? So every data center so that's wasted. I'm not talking about the total power draw just the wasted. So there is there are so many dimensions to this um to this problem of we are not yet really good at efficiently operating the underlying infrastructure and extracting maximum utilization. In a nutshell that is the problem that we fundamentally attack. If you think about sort of where does clockwork intersect that problem, Chris, we so AI workloads are they achieve everything they do through a distributed communication process. A single GPU cannot do what you need done.
47:53 So you're throwing thousands of GPUs to complete the job. And so really the bottleneck to how you accomplish whatever your workload is trying to do is all about communication efficiency. And and that's why sort of we we we we like saying that communication is the new Moore's law in AI infrastructure. That's because everything that you're trying to achieve from an AI workload depends on the efficiency, effectiveness, reliability of how these GPU GPUs talk to each other in a GPU cluster.
48:22 >> Yeah. Yeah. Well, I mean, and and I've heard it des your company described as sort of the ways for GPU clusters. I mean, so that's because everything that you're trying to achieve from an AI workload depends on the efficiency, effectiveness, reliability of how these GPU GPUs talk to each other in a GPU cluster. It's it's helpful to start with sort of what are the three building block technologies, Chris, and we've touched on a couple of those, right? So, the first building block is is uh and I'll talk through there are three core building blocks. So, let me walk through each one of those. The first one we've touched on is already our ability to synchronize clocks to within tens of nanconds. And as a result, sort of what we are able to do is to look at every single message. Um, so you take a pytor job, you break that down into sort of what is a distributed communication. So there's a messages being sent by this collective communication library. That message breaks up into chunks. Those chunks are broken up into uh units that are sent over RDMA connections. And so you break it all the way from the application down to what's traveling on the wire and you're timing everything that's traveling on the wire because of our ability to synchronize clocks very accurately. So the first thing we have is the ability to look at all communication understand what's going normally and what's being delayed map that to what does it mean for your pyarch training jobs if you will. So that's the first building block is clock sync leading to really deep network telemetry correlated to the application.
49:54 The second building block is what we term dynamic traffic control. So really it comes to uh it's actually very simple. What we are able to do is pace traffic. So either slow it down when it's likely to be congested. So pace traffic we can slice traffic. Um, so we basically are able to say if I want to slice it into five pieces and guarantee one versus the other, we can slice traffic and then we can reroute traffic, right? And so that's really what this control plane does. We happen to plug it into TCP using instrumentation like EVPF instrumentation. We plug it into RDMA networks using RDMA APIs. We plug it into collective libraries like Nickel within Nvidia or Rickle in the AMD ecosystem. So this ability to slice, pace and route traffic works with different communication protocols, libraries and APIs. That's the second foundational capability. And then the third one is what we call distributed state tracking. So sorry if I'm hopefully this is making sense and I'll I'll translate this but because all these jobs are essentially being managed as a set of processes that are executed on large number of GPUs. Imagine that one GPU fails. If you're not making sure that everybody is at a consistent point, then essentially you have to restart from the scratch because where you failed the job is hung and all the GPUs are at a different place if you will in that in that collective process. And so what we are able to do extremely well is understand what the distributed state looks like so that we can recover gracefully when failures happen when links. So that's really the third aspect of it is understanding the distributed state in a in a job that's executing on large number of GPUs and that allows us resilience and fault tolerance when failures occur. So if I step back we at the core it's three breakthrough technologies right one is clock sync leading to insane telemetry. The second is dynamic traffic control that allows us to sort of pace, slice and route traffic flows. And the third one is distributed state tracking that allows us fault tolerance if you will. And so that's what we've combined. Now if you think about what our customers see when they're operating these clusters, we have um multiple large NeoClouds that are now deploying us. Uh these are among the world's most successful NeoClouds.
52:19 We have one of the world's leading um well LinkedIn was on our uh uh uh on our webinar yesterday. We have one of the world's leading telecom companies uh deploying us, leading uh sovereign labs and one of the top five hyperscalers, right? So good deployments that give us a broad sense of um what we're seeing and essentially it's three things. One is observability. So we first get deployed because we give extremely good observability particularly with respect to networking but broad cluster availability. In fact, one of our Neoclouds uses us to audit their cluster to make sure it's configured correctly before handing off to their customers.
52:57 So, so really good observability of the cluster. Then, um the second and arguably the most common reason why we get drawn in or the biggest deal driver is resilience. Our ability to keep a training job continuing without disruption when you have links flapping. Today we're working on technologies that will allow us to survive not just a network link flap but any kind of GPU failure without transmitting that failure up to the application and continuing non-disruptively. So fault tolerance is the second thing that we're deployed for. And then the third one is performance optimization. So just by using sort of load balancing congestion control how do we optimize the performance of the network and therefore the overall cluster. So that in a nutshell is sort of what what we do.
53:44 >> Yeah. And I got to imagine the TCO on that is so easy >> too. >> It is. I mean honestly I think it's uh if you just take the link flapping u this is sort of our analysis right. So u and and this is based on really good data both published data from the likes of Alibaba and others that have really meta and others have documented failures in gory detail but we've also seen this in our own customer base.
54:12 For every thousand GPUs, you're likely to witness something like 150 to 300 failures a year. Uh failures in the sense that restarts, job restarts a year. And then you continue and say, let me do the math on those restarts. You're wasting somewhere around 200,000 to 300,000 GPU hours a year from link flaps alone in in a GPR. So 200 to 300,000 GPU hours. So you can do the math on just that. uh forget delayed um time to production and all of that just on the lost hours at $2 to $2 and a half dollars per GPU hour. The math there itself even on a 1,000 GPU hours you're getting to sort of close to a million and and that doesn't count the human cost of just finding problems. It doesn't count the delayed time to market of your products and so on and so forth.
54:59 So I think the ROI has been uh easy enough to establish for for sort of the value uh Chris >> for sure. Hey, I just just for a second. Um I'm noticing like the the light behind you >> from the open window is like shifting and it's like it's picked up so it's it's got the little >> Oh, yay. Can't see your face. >> Yeah, it's it's starting to change the exposure bit. Yeah, that's that's probably a better bet.
55:26 >> Perfect. Perfect for that. >> Yeah. Yeah. No, it was just getting a little little uh blown out there. No question you said something, Chris. Thanks for saying here. >> Yeah, just so you know, >> it started out all right, but I think the sun, while we were talking, the sun moved and the clouds changed. >> Yeah. Yeah. >> Chris is the pro at this. I don't know if you know, but he's like, he's on multiple podcasts. So, >> uh, this is not his first rodeo.
55:52 >> I know. His setup was impressive last time we saw it. So, yeah. [laughter] >> What do you do? I do four podcasts now, so crazy. >> Yeah. Told you. Nerd talk. >> Nerd talk. >> Nerd talk. Oh, god. That's Chris's jam jam. Uh, one of the the key takeaways for me and anyone that is watching, you know, from an infrastructure perspective, there's a rejuvenation that is happening. However, I'm not seeing the average uh, you know, Fortune 500 company building out these very large scale environments. The problem you are solving just by its nature, you need thousands of GPUs. That's right. So if you're spending thousands, you know, MI billions of dollars, it certainly makes your value proposition uh um you know uh very simple. But what my biggest takeaway was I didn't realize how many failures there actually are >> indeed. Not just the just the the GPU itself, but also the network. It and I just kind of felt like why doesn't this just work better, you know? No, I I I think it will it will take a few years.
57:01 I think the ideas around how to make the infrastructure perform more resiliently to have more observability built in to have more resilience built in to have um security built in will all come I think we're at a phase in the evolution of AI infrastructure deployment of GPU clusters where speed trounces everything and so everybody's going for larger data centers faster and so there's not enough time to step back and design and and the other phenomenon in here. That's interesting. I want to come back to the question about enterprises and how many can actually deploy these versus how they'll consume AI. Will they consume it through the infrastructure layer or do they basically consume it through some other cloud layer? Um, but sort of what I was going to say is many of the people that are solving these problems are solving it in a bespoke manner for their own internal infrastructure. These are the open AIs of the world and the anthropics of the world are likely solving these problems through software techniques that are specific to their needs.
58:03 >> What you've not yet seen is the emergence of independent software companies that are starting to say all of these are software value ad on top of the underlying networking hardware on top of the underlying GPU hardware server hardware that we need to build out that's existed. I mean there isn't a data dog for observability in the GPU world. there is so there are many many software ISVS that are just I think going to emerge over the next few years and that's certainly our vision >> you asked a great question right so will I mean this is something I think about all the time given the complexity and the cost of this infrastructure I think there won't be there will be always sovereign AI uh that will be sort of deployed uh as bespoke data centers but for and there'll be some really large enterprises that have the scale to justify building out their own AI data centers. For the most part, I think the emergence of the NeoCloud space is because you want someone else to take these problems away from you. And I think um on the one hand going all the way to an AWS and a Google as they're becoming more and more AIcentric, it's it's sort of it's almost like I don't need to consume their 150 cloud services or 250 cloud services. I just want a set of services tuned for AI. Um and that's where the neo clouds came in. They are really focused and therefore lower cost and uh more purpose-built for AI.
59:30 Gradually you're seeing those offerings emerge from the big hyperscalers as well. So I think most enterprises over time will consume that from either hypers scale public clouds or neoclouds that are becoming larger and larger and almost um challenging the hyperscalers. And so uh there will be uh some really large enterprises of course that will do their own but that's so for us part of what I'm excited about when I see the success we are having with um the hyperscaler I mentioned and the and the neoclouds ultimately we are serving the tenants of these companies we also some of our products are more useful so to the tenants than they are to the operator themselves right so for example we do network monitoring as I mentioned and there's a lot of telemetry at that layer, but we also expose infrastructure dependencies all the way through a PyTorch lens. So if you're an ML team that wants to know why is my iteration running slow and what can I do to change that, we give you insights on how to correlate what you can change at the infrastructure layer all the way through to your training jobs. Interesting.
60:34 Similarly, if you think about having a PyTorch job continue undisrupted when failures happen, what we're hooking into are hooks within nickel hooks within PyTorch itself that are relevant to the tenant of these clouds whether you're running in a Google or an Azure or you're running it in a Nebus or in your own data centers. >> Yeah. Well, you know, it's an interesting conversation that's been happening around the idea of like have we with all these foundational models, have we sort of meet reached peak uh training and now it's all going to be inferencing. Yeah.
61:10 >> And then you you start >> then you start looking at like well but you know we've still got all this visual training data that's going to be coming in and there's just like this is going to happen, this is going to happen. It seems like training is not going anywhere anytime soon in my book. Um but like what's your perspective on that? >> Yeah. So I think um for sure >> uh training I think is nowhere near the point where you can say there'll be sort of no need to train and everything will work on a foundation model. I do think that it's hard for me to imagine more than a handful of foundation models that will be successful over time. So if you think about sort of truly massive foundation models, you'll maybe have less than 10 globally. I mean there'll be some that countries will want to have simply for national security reasons and so on but in terms of broadly applicable foundation models but then uh let's take a simple example I mean actually you you called out image models as a great example there will be a few uh dozen companies really specializing on imagecentric models um you take um automotive um Tesla is at least publicly known to have deployed somewhere between 50 to 100,000 GPUs just for training automotive and it's not likely that there's going to be a single model that folkswagen and Mercedes and everybody else will use. So the automotives there will be a couple of dozen uh self autonomous driving models in uh that will exist. Similarly, drug research in pharma, you'll continue to have either not necessarily groundup models, but pre-training and fine-tuning and there's a whole bunch of and so when I think about it that way, I'm convinced that there will be several hundred large companies doing training.
62:56 >> I don't believe Chris that there will be tens of thousands of enterprises doing training. they may do some amount of sort of fine-tuning but even that I believe is is probably not going to be numbering in the tens of thousands inference of course is frankly every application in the world will be an inference application over right and so so I think that is truly extremely broad >> the the the thing interesting thing about training is if there are a thousand companies doing training the infrastructure spend there is still enormous enormous >> what I also find interesting is that inference itself is becoming multi-GPU highly distributed.
63:35 >> Yeah. >> Um there are trends that are driving sort of as as you have larger and larger context length and keeping the what's called KV cache uh which is really the tensor translation of all the context if you will that needs to be in memory to produce the output when someone asks query. the amount of um context whose tensor values you're storing up front uh and to to invoke rapidly when a new query comes in that's growing into the terabytes of memory and that's forcing a disagregation of inference into multi-GPU inference if you will >> and so that is bringing >> that sounds like a perfect application for this product I know namely [laughter] exactly no exactly we are finding that suddenly all the things we talked about as highly stateful, highly latency sensitive uh multi-GPU training environments are becoming more and more true of inference as well and that's what I believe will happen even in the world of inference.
64:39 >> Yeah, I I think right now the the biggest challenge for AI is is context. >> Yes, exactly. Exactly. Exactly. And there are some really powerful technologies that are evolving in the infants world to solve that. Right. And so how do I first is how do I store hundreds of gigabytes to terabytes of inference in an external storage system and then and yet bring that into the memory of a GPU that's actually serving that inference request in a really fast manner using special um u protocols to move data fast for inference serving.
65:13 How do I take an in incoming query and route it to the specific GPU that has the context most readily available and is idle at the same time? Right? And so all of it's a it's a it's a it's slowly a communication optimization problem that that you're starting to see for sure. >> Yeah. Yeah. What's interesting is you're when new technologies come out, it's rare that they start immediately in in these large, highly performant, complex environments, >> right? And and that's that's what's interesting to me too is that you're >> if you can the customers that you have, if you can help them and you can solve their problem, >> the others will be much easier. you have gone after the hardest customers out there and you're trying to sell them.
66:05 >> No, that's so true. Um, it's it's you know, it's it's both the the boon and the bane of of uh the problem we've solved and where that problem is most urgent. um even in our non-GPU microservices um customer base where we are optimizing um sort of both detection of network bottlenecks in large scale microservices and addressing the problem. the customer that I think we in our public announcement of the of our plat of our coming out party if you will we talked about Uber as one of the customers and what's publicly known about Uber is they have thousands of microservices running on nearly 200,000 machines across three clouds right and so so so the scale where getting latency down right uh is matters the most is in these extremely large scale userfacing applications where a delay or an outage on an application means revenue is on the line, right? Rides are lost or food is not delivered. And so there's real implications to service levels there. On the GPU side, in a similar vein where our ability to deliver resilience matters is when you're running thousands to tens of thousands of GPUs and when you lose time, it's basically thousands of GPUs lying idle. So the so you're going to do whatever it takes to fix that. And so the good news is we are battle tested in some really large environment and the revenue per customer is abnormally high compared to all my prior experiences. The bad news is failures are very very I mean you your product had better do what it's saying because otherwise the black eye is is dangerous.
67:46 >> Yeah. No doubt >> to a startup in particular. Yeah. >> Yeah. Yeah. I mean I I could keep going for hours. Um we'll try to land this plane. Um, we have already just blown through our time. So, >> no, I apologize and thanks. >> And I'll tell you, man, the the the the first problem you solved is a really interesting one, too. And I I feel like we almost need a podcast just to talk about that technology.
68:11 >> Yeah. To to get it down to that level of granularity. >> Yeah. >> Is really interesting. I I'd be very curious. >> Chris, this is the honest truth. when when I first um even though I knew the founder for 15 years our like my my interaction with clockwork started when a uh head hunter called me and described this company that had solved this clock sync problem and um I my first reaction was look I've been I've been in tech for a long time this is a very old problem in computer science how do you synchronize and frankly if blocks could be synchronized to that level of accuracy databases would not ques s in order to do sort of right in order to do some housekeeping. Storage systems would not qu in order to take a snapshot. They would just rely on timestamps to say I can create a snapshot and so on. So I'm like I'm not sure I buy what you're saying. And so uh and in fact he then sent me the paper um the Usenix paper where this was presented. And that's when I realized oh my gosh this is a this these are people I know well so let me just go. I didn't know that's that's what they were working on. So I agree with you. It's an extremely uh interesting problem >> and that it's all in software. That's the thing that really exactly gets me because I'm not I'm like trying to envision how you would do that >> and no you're you're of course right because this problem has been solved through hardware. PTP has solved this to the same levels of accuracy but with very specific hardware on a smaller number of machines. Exactly. Exactly.
69:39 Exactly. >> Exactly. >> Super interesting. So cool. Um, well, landing the plane. Um, are you before I ask you this question, I just want to know, is it okay to talk about torch pass? [clears throat] >> Yeah. No, I I alluded to it by saying we're now surviving other GPU failures. That's what we're working on. We're actually in I'll te you up for that. I just want to make sure that was okay. I don't know where you were at announcing it.
70:04 >> So, Sur uh what is in store for Clockwork now both from a technology vision and strategy as well as the company? >> Yeah. So I think um I'll start with the technology. I think our our the vision we have of a software-driven fabric has so many places where I think we can add value. I'll call out two initiatives in particular Sesh. The first one I talked about our current product capability to survive network link flaps and keep training jobs running. We're working on a project um internally we call it touchp pass but basically what we're what what it's aimed at is the ability to survive any kind of GPU failure and continue to uh operate training without letting that failure disrupt Python jobs if you will. So non-disruptive fault tolerant training in the face of any kind of GPU failures not just network link flaps is one of those um projects.
71:01 The second one is uh really being able to allow RDMA storage and TCP flows to coexist on a single Ethernet fabric. Um specifically RDMA requires a very separate set of configurations on underlying Ethernet even if you use Ethernet. Rocky Rocky V2 is the pro is is how you deploy RDMA on Ethernet fabrics. And then TCP of course is extremely wellnown. Traditionally the two behave extremely differently. TCP uh anticipates losses and uses losses to control congestion. RDMA requires lossless networks and so typically to deploy them you've had to physically segregate the network uh in in very complex ways. Our software very trivially allows you to run both RDMA I don't say trivially I should say in a very simple manner allows you to run RDMA and TCP on a common underlying Ethernet fabric. We're very excited about this. It's early stages. We have one of the largest hyperscalers triing this in their environment but it can lead to massive cost reduction and simplicity. So those are a couple of but the larger vision of ultimately allowing software to write to really uh optimize distributed applications on commodity networks. So these are examples of what we're doing next that go in that direction that I'm excited about. As a company to be honest, what's top of mind for us is just scaling the business both in terms of customers but really scaling the employee base as well. We're based out of it's a young company. We've uh added almost uh we've grown by almost 40% in the last few months and sort of scaling the engineering team, scaling our ability to deliver to really large customers is is top of mind. Uh and so that's sort of something else I'm excited about. It's it's a >> it's fun at that stage indeed. It's very challenging cuz it's finding good talent uh especially given a strong desire for us to sort of be in one location, have everybody sort of literally work from a single office and so on is is something that we're we want to hold on to as long as we can and so it makes it even more challenging, >> right?
73:09 >> Yeah. Yeah. Well, you you seem to like challenges, so um I I I have a feeling you'll you'll figure this one out. >> I'm not sure you're going to make it to retirement. I think that [laughter] just as long as life is fun, why bother retiring? Exactly. >> Exactly. >> This is retirement for you. >> Indeed. Indeed. Indeed. Getting paid to do what >> interesting too, you know, from a sales guy perspective. >> You you have It's not like there's thousands of potential clients for you.
73:37 You're really focused on the big big guys. But also, this is a very technical sale. >> It seems like you really need to understand the these components and how they're talking to each other and you know all these resiliency issues and the visibility issues just >> there's so much that goes into this >> so I think it's interesting >> you know what I mean like you can't just have an average but I I'll say one thing sesh um I'll come to the first point second but let me start with the second point of it's a very technical sale I'll say yes and no because I think on the on the yes side when it comes to the how do you solve the problem and you explain how your technology works. You definitely need a strong technologist on the other side that can that can first buy into how you're approaching the solution space and then prove it to themselves in a POV. There is no deploying our solution without testing it in your own environment and so on. So that but the early phase of can I get to that second technical conversation and a proof of concept what we are finding is articulating the problems >> and asserting that these are what you're probably seeing in your environment the person on the other side quickly says exactly right and so the problems are very easy to articulate they are something that they're experiencing all day long so getting our sales teams are able to go to the right as long as you go to right person and say this is probably what you're experiencing. If so, we want to have a second conversation that's technical. That's proving to be not that hard. And so, at least so far, um I think Joe and the sales team is is generally having a good time of being able to engage customers and get to the next conversation.
75:21 >> Yeah, it's a wellrecoognized problem. >> The financial sale is what really makes me interested, you know, I guess you're right. Yeah. So like um yeah you need a technical person to explain it but the business value is really clear. >> Exactly. Exactly. And and you know Joe is a big Joe our CRO Joe is very big on saying look I need to be able to translate this into what it means to the customer and their business. And he's saying this is not hard for me to do in this case. And so that's step on the question of how large is the customer base. Um something else I'm excited about is is um so today we focus on really two very targeted customers. One is neoclouds. There's about 180 I believe of them and and sort of >> uh over time it'll probably become a smaller number but that's one group. And the second one is extremely large enterprises that have more than 256 GPUs that are doing training. Now in the not tooistant future we are bringing something that allows you to address people doing training in clouds uh with a small number of GPUs in let's say an Azure or a Google and bring some of the value even to those environments and so that certainly expands the group to a much larger audience of who we can solve. Part of it is also being very careful to say let's frontload where our sales and marketing efficiency can be uh highest where your sort of um large deals with technically sophisticated buyers is where we can make the fastest progress and so we'll stay focused on the first audience but gradually we're seeing an opening up of the opportunity to even sort of cloud hosted training customers and as I mentioned inference is feeling very much like a distributed application and so that's the next uh that will open it up larger. So we see a pathway today for sure it's focused on 500 to,000 customers as the as the focus.
77:12 >> But what a great place to start because it's such a high target. >> Yeah. >> Exactly. I worry more about branching out sooner than we should than about not branching out. >> Yeah. Yeah. Well, and I and I got to imagine as as you evolve over time, there's some foundational technologies that you guys have developed there >> be used in a lot of different areas. I could not agree more, Chris. I could not agree more.
77:38 >> So, we're going to have to stay in touch. Surish. Um, >> absolutely looking forward to that Chris. >> So exciting that where you're at in such a short period of time here and just this conversation I and learning about clockwork. Um, I'm learning so much. Uh, I learned so much just through the conversations I had with Joe and you and just reading about you guys. I joined that webinar as well. Um, so hopefully you're a frequent offender on here and uh, >> look forward to it and yeah, we're all living in fun times again. So, for >> sure. Awesome. Well, thank you so much for joining.
78:08 >> Thank you. Take care.
Summary
- The resurgence of infrastructure skills is critical as AI and GPU workloads demand high-performance computing capabilities.
- Clockwork is developing technologies to enhance GPU cluster performance, focusing on clock synchronization, dynamic traffic control, and distributed state tracking.
- Current GPU clusters often experience low utilization rates (30-50%) and significant downtime due to failures, leading to substantial financial losses.
- The conversation highlights the importance of observability and resilience in AI infrastructure, with Clockwork's solutions addressing these challenges effectively.
- There is a growing trend of enterprises relying on NeoClouds and hyperscalers for AI infrastructure rather than building their own.
- The discussion touches on the potential for new software solutions to emerge that optimize AI workloads on existing hardware.
- Future innovations will focus on enabling non-disruptive training in the face of GPU failures and allowing RDMA and TCP to coexist on the same network fabric.
- The speakers emphasize the need for clear communication of business value alongside technical solutions to engage potential customers effectively.