Section Insights
Introduction to Flappy Airplanes
Who are the founders of Flappy Airplanes and what is their background?
The founders, Ben and Asher, have backgrounds in GPU systems and machine learning. They clarify that Flappy Airplanes is an AI lab, not an aviation company.
- Founders have strong technical backgrounds in GPU systems and AI.
- Flappy Airplanes focuses on data efficiency in AI, not aviation.
- The company is newly launched and has received interest from various industries.
The Importance of Data Efficiency
Why is data efficiency crucial for AI capabilities?
Data efficiency is essential because current large language models (LLMs) require vast amounts of data to perform well, while humans can achieve similar results with significantly less data. This efficiency is vital for expanding AI applications in areas with limited data.
- LLMs excel in tasks like search and coding due to abundant data.
- Achieving similar capabilities with less data could unlock new AI applications.
- Data efficiency is key for sectors like robotics and scientific research.
Challenges in Data Collection
What challenges exist in collecting high-quality data for AI?
Collecting frontier-quality data is complicated due to the lack of centralized data providers and the need to navigate regulations and business negotiations. This complexity makes data-efficient models easier to deploy.
- Data collection for AI is hindered by regulatory and logistical challenges.
- A more data-efficient model could simplify deployment significantly.
- Centralization of data limits the ability of companies to train AI models.
Innovating with Hardware and Frameworks
How does Flappy Airplanes approach hardware and algorithm development?
Flappy Airplanes aims to explore new capabilities by optimizing existing hardware, particularly GPUs, and developing new algorithms that current frameworks struggle to express efficiently.
- Flappy Airplanes focuses on maximizing GPU capabilities for AI.
- Innovating within existing hardware can lead to significant advancements in AI.
- The company seeks to push the boundaries of current machine learning frameworks.
New Systems and Algorithms
What is the significance of building new systems for AI?
Developing new systems allows for the creation of algorithms that can address data efficiency challenges. These systems enable innovative approaches to AI that are not possible with current frameworks.
- New AI systems can facilitate the development of novel algorithms.
- Creative approaches to AI can lead to breakthroughs in data efficiency.
- Collaboration with diverse backgrounds can drive innovation in AI.
Transcript
0:03 Bringing up two frontier spotlights, the first is Ben and Asher, two brilliant brothers. Yes, they are come on up. founders of swimming submarines. Wait, no. Wait a second, flapping airplanes. and we're excited to hear about data efficiency. Thanks so much. Thanks, Konstantin. >> >> All right, yeah. It's really great to be here. Thanks for having us. we're very excited to tell you about why the future is data efficient. so, let's get into it. So, first I think introductions are in order. my background, I spent 3 years in the PhD kind of deep in the kernel mines writing very low-level GPU systems before helping to start Flappy. And I I also should say I I at one point helped start an incubator called Prod that worked with a bunch of companies that did well. So, that's also part of my background.
0:52 Cool. And so, I'm Asher. if you can't tell, I'm Ben's older, wiser, and slightly less handsome brother. I previously was also a Stanford PhD, spent time at Cursor, Mercor. And then lastly, our third co-founder is Aiden Smith. he's a Thiel fellow. He's busy working right now. but he's spent 3 years super commuting between his college and our lab. So, he knows a lot both about the brain and also about machine learning. so, before we get started in earnest, I do want to clarify some confusion. We launched the company 3 months ago. we are very excited to have received a lot of inbound, but a surprising amount of it has been from the aviation industry.
1:26 you know, people trying to sell us things like runways, like not not runway in the venture capital sense, like literal runways. you know, airplane parts, wind tunnels, etc. So, I'd like to disclaim once and for all, we are not an airplane, we're an AI lab, and hopefully that will become clear by the end of this talk. Okay, great. so, first a quick outline. In this talk, we're going to tell you about two things. First of all, our thesis of why this is a really important problem. And second, we'll tell you a bit about our approach that fuses both systems work and algorithmic work. So, first on the thesis side.
1:57 So, if we look at the current state of the world, LLMs have gotten incredibly good at a couple of really valuable tasks. So, for example, search, coding. Together, these are, you know, at least a trillion-dollar market that's probably accessible here. A big part of the reason why they've gotten so good at these tasks, and this is related to what Andre I think talked about this morning, is that these are incredibly well-resourced tasks in terms of the amount of data that's available for them. For search, this is basically the entire internet. For coding, this is a big fraction of the internet. Coding is also a very friendly environment in that it's very easy to produce mountains of synthetic data if you want to.
2:26 so, when we talk about data efficiency, what we're asking the question of is, is it possible to get these kinds of capabilities with much less data? Intuitively, it seems like it should be possible. you know, humans are able to become quite good at coding with, you know, maybe 10,000 times or 100,000 times less data on this than these current models take. So, this is kind of the core question. And a reason why this matters is that if we look towards the future, there's a bunch of other domains of the economy where this where there's there's much less data. So, for example, you know, Jim talked this morning about how in robotics, there's a ton of work that has to go on to try to generate the data.
3:00 It's much more complicated than it is in in coding and search. In trading, there's only so much financial data. Scientific discoveries, obviously very little data. but the the the potential there is sort of unbounded. And most important of all is, of course, the end-to-end toaster supply chain. Now, the reason I bring this up is not because of what it is, it's because of what it represents, namely that there are tens of thousands of things like this that really make up the broad economy. And the broad economy doesn't just look like search and coding, it looks like a lot of things that are actually really not well-resourced.
3:28 Yeah, and so Ben's given you one economic reason to believe that the future is data efficient. I think a second claim is that compute is actually easier to scale than data. I mean, we know that flops get exponentially cheaper over time. And you know, I think it's probably true that data is not getting cheap as quickly as compute is. Although, it is getting cheaper. I think, you know, the second reason is that collecting frontier quality is complicated. The the compute market is homogeneous in a or is more homogeneous than the the data market is. Like, you know, Greg was telling us after, you know, GPT launched, they could just try to buy all the compute. And there's no centralized data provider. Like, if you want to go into the economy and collect the long tail of tasks, you have to like deal with regulations, you have to negotiate with businesses about terms of use. It's actually very annoying.
4:08 So, I think if you could As a result, if you can make a model that's a thousand times more data efficient, I think it'd be a thousand times easier to deploy. So, those are two sort of economic reasons that we think data efficiency is important. The last reason is actually kind of philosophical. I think that, you know, if you look at the world today, there's not that many companies that can train AI models. And And you know, part of this is because of of centralization of compute. But I think it's also because of centralization of data. Like, you know, I've heard news about Neo Labs who, when trying to create new capabilities, they like literally buy out distressed bookstores to find all the the data they can find.
4:43 They like go to like rare libraries to find all the tiny niches of data that you need to actually make a frontier model. And I think in a world that's data efficient, I think that companies can actually more broadly participate in the AI revolution. Like, you know, we heard earlier today that you know, in in the poll about what the moat is in the AI world, the the most common answer was data. If that's true, then data efficiency is the thing that actually enables broader competition. So, I think if you care about the shape of the world to come, I think you really should care about data efficiency because it it modulates who can actually participate in which parts of the AI economy.
5:15 Cool. So, now we're going to talk about our approach. so, our goal is to design data efficient AI. we design algorithms for this. We're not going to talk about them cuz they're our core IP. But I will tell you about an important facet of our approach, which is how we try to look in new places. Like, if you want new capabilities, where can you find them? So, our claim is that essentially, if you want to develop new algorithms, you should look at new primitives for interacting with hardware. There's some set of things that GPUs can do efficiently.
5:41 And then there's a smaller set of things that current frameworks like PyTorch, for example, can actually express efficiently. And, you know, that that gray circle, we've seen a lot of research into that circle. We actually know what kinds of algorithms work. But if you're looking for new capabilities, you might want to look in these new places. And that's where we live at Flappy Airplanes. That's what we try to do. And you know, Ben will talk for a little bit later, for example, about things like fine-grainedness, which actually are hard to do under current frameworks, but a GPU actually can do.
6:08 you know, history suggests that this is a good bet, for what it's worth. I mean, a lot of the development of machine learning over the past 15, 20, even 100 years has actually been new primitives for interacting with hardware. Not necessarily designing new chips, which is also important, but merely figuring out how to squeeze more out of the existing technology that has proliferated. Yeah, so I guess a bit of context. in my PhD, I did some of this work on on early work on mega kernels, which was trying to make GPUs do weird stuff.
6:35 at Flappy, I would say that we are going further in spiritually similar directions, namely trying to kind of abuse the crap out of GPUs in ways that haven't been done before. this has been quite a lot of fun. So, I actually would recommend it if you if you like playing with hardware. And I guess part of what I'm excited about in this talk is that I, you know, this is in general I think a fairly high-level venue, but I actually want to get into the weeds of some stuff because I think it's fun.
6:58 So, if we look at how kind of current systems work, and this will be a little bit abstract, but I hope that the the concept will come across. What makes current machine learning frameworks easy to use is that they synthesize a single-threaded programming model out of very parallel processors. So, basically, when you write PyTorch, you write matmul, and then attention, and then matmul, and then RMS norm. And under the hood, there's all kinds of contortions that happen where the software dispatches it in parallel across a very parallel set of processors. But nonetheless, to the programmer, this looks very natural.
7:26 But what happens if you want to run things that look like this? Or maybe you want to run things that look like this. All of a sudden, these kinds of things are not easily expressed in current frameworks. So, one thing I'm just very excited to show sort of a little teaser of is our internal framework that we use, which is built off a virtual machine that kind of just takes over the whole GPU, and we just do everything we want ourselves on it. in this particular case, this is not a real workload I'm showing. This is kind of a stylized thing that is similar, but gives is an example of something that is also asymptotically inefficient to run in PyTorch. This is a very small batch kind of deeply pipelined hogwild-style training loop. this is the kind of thing that's very hard to do with with current frameworks. So, the the thing that I sort of like to get across here is that when you build these new kinds of systems, you enable new algorithms. Many of these algorithms we think are actually very relevant to the data efficiency problems. This is why we care about it.
8:21 and that it is the kind of co-optimization of those things that we find really interesting. Cool. So, we're almost out of time. I I just really have one more thing to say, which is that if you find this exciting, please come talk to us. you know, we we work with lots of folks who have trained big models before, but we're also trying to do creative new things. We're trying to change the paradigm. And you know, certainly our experience watching some of the the companies Ben has has helped incubate is that, you know, creative people with unconventional backgrounds can actually do amazing things and change the world.
8:47 So, if you are of that background, please come chat with us. Okay, so I guess maybe a very quick recap, and then I think we probably have time for one question or so. first of all, data efficiency. If you care about the shape of the world to come, if you care about the broad deployment of AI into the economy, you should care about making models much more data efficient. And that you know, our finding has been that building new systems to enable more fine-grained use of GPUs is actually very relevant to building these kinds of systems. and as Asher said, if you find this exciting, please come talk to us for a round.
9:19 Thanks, everyone. Thank you. >>
Summary
- The future of AI relies on data efficiency, allowing capabilities with significantly less data than current models require.
- Many economic sectors, such as robotics and scientific research, have limited data availability, highlighting the need for data-efficient AI solutions.
- Scaling compute resources is easier than collecting diverse data, making data efficiency a strategic advantage.
- A more data-efficient model could democratize AI development, enabling more companies to participate in the AI landscape.
- Flappy's approach combines systems work and algorithmic innovation to enhance data efficiency.
- The founders emphasize exploring new hardware interaction primitives to unlock new algorithms and capabilities.
- They advocate for creative and unconventional approaches in AI development to drive innovation.
Questions Answered
Who are the founders of Flappy Airplanes and what is their background?
The founders, Ben and Asher, have backgrounds in GPU systems and machine learning. They clarify that Flappy Airplanes is an AI lab, not an aviation company.
Why is data efficiency crucial for AI capabilities?
Data efficiency is essential because current large language models (LLMs) require vast amounts of data to perform well, while humans can achieve similar results with significantly less data. This efficiency is vital for expanding AI applications in areas with limited data.
What challenges exist in collecting high-quality data for AI?
Collecting frontier-quality data is complicated due to the lack of centralized data providers and the need to navigate regulations and business negotiations. This complexity makes data-efficient models easier to deploy.
How does Flappy Airplanes approach hardware and algorithm development?
Flappy Airplanes aims to explore new capabilities by optimizing existing hardware, particularly GPUs, and developing new algorithms that current frameworks struggle to express efficiently.
What is the significance of building new systems for AI?
Developing new systems allows for the creation of algorithms that can address data efficiency challenges. These systems enable innovative approaches to AI that are not possible with current frameworks.