Section Insights
Introduction to Atlas and Spatial Intelligence
What is Atlas and how does it contribute to spatial intelligence?
Atlas represents a significant advancement in spatial intelligence by enabling the generation of spatially contextualized pixels. Unlike traditional models that rely on extensive camera setups, Atlas can achieve similar results with minimal equipment, thus unlocking new possibilities in video generation.
- Atlas uses next view prediction to enhance spatial intelligence.
- It can generate high-quality video with fewer cameras than traditional methods.
- The model's performance improves with increased size and training duration.
Understanding Spatial Intelligence
What is spatial intelligence and why is it important?
Spatial intelligence involves generating, reasoning about, and interacting with space, which includes understanding geometry, structure, and physics. Atlas makes strides in this area by providing crucial information about camera viewpoints, essential for simulating and rendering environments.
- Spatial intelligence encompasses 3D and 4D capabilities.
- Understanding geometry and physics is fundamental for effective spatial reasoning.
- Atlas's advancements represent a major step in achieving true spatial intelligence.
The Interplay of Generation and Reconstruction
How do generation and reconstruction work together in spatial models?
Generation and reconstruction are interdependent processes in spatial modeling. While reconstruction requires multiple viewpoints to triangulate 3D points, generation fills in gaps where data is missing, ensuring a complete representation of the environment.
- Reconstruction needs multiple views to accurately represent 3D space.
- Generative processes are essential for filling in unseen areas in 3D models.
- Even expert captures can miss details, necessitating generative capabilities.
Persistent 3D State in Creative Processes
Why is a persistent 3D state important in creative applications?
A persistent 3D state allows creators to build and manipulate environments over time, which is crucial for industries like film and gaming. Atlas aims to provide stability and consistency in 3D modeling, enabling users to create coherent and evolving spatial narratives.
- Creatives require stable 3D environments for effective storytelling.
- Atlas seeks to enhance control and precision in spatial modeling.
- The ability to maintain a persistent state is key for spatial reasoning.
Future Directions and Dynamics in Spatial Models
What are the future directions for Atlas and the importance of dynamics?
Future developments for Atlas will focus on incorporating more dynamic elements into spatial models. While current capabilities include basic dynamics, enhancing this aspect is crucial for applications in robotics and other fields where understanding movement is essential.
- Incorporating dynamics is vital for realistic simulations.
- Atlas currently has foundational dynamics that can be expanded.
- Feedback suggests a need for more dynamic interactions in future models.
Transcript
0:00 On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken. >> You know, LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction. >> This is the real place where AI can actually unlock a ton of value for people and their process. We're saying like 50, 100 extra reduction. There's a famous shot in the first Matrix movie where Neo is like falling down. Exactly.
0:27 They had hundreds of cameras doing that angle on a green screen. With Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration. >> No one has ever seen these results. >> When you set out to do this, did you know it was going to work? >> I was pretty sure. Each time we made the model bigger, and each time we trained it for longer, it got significantly better. >> Does that mean we're going to get 4D video? Can I go walk around?
0:51 >> So, big day yesterday, you launched a new frontier model, which has got amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. So, maybe Justin, do you want to talk about what was launched yesterday, why it's significant. >> Yeah, so Atlas is our new next generation world model. it has three basic things. It can generate, reconstruct, and simulate the world.
1:12 so within that, there's a couple different major capabilities. It has really good camera condition generation. So, you can input an image together with a camera trajectory with a camera trajectory and steer the model and have it generate, you know, video frames along any any perspective you want. It's really good at sparse 3D reconstruction. You can input one or multiple up to 100 frames, that are views of the real world, and use those to reconstruct the real world. And that reconstruction can take the case either of of novel of video flying through the space, or an explicit 3D reconstruction of the space.
1:41 then finally, it can be used for simulation. and for this, we show off, you know, these awesome bullet time videos, which got a lot of attention online, and then also robotics simulation. >> What's a What's a bullet time video? >> A bullet time video, this comes from the Matrix. You know, there there's a famous shot in the first Matrix movie where Neo is like falling down. >> >> Oh, yeah. >> Exactly. So, and then remember in that famous shot he's like falling down, it's in slow motion, and the camera flies all the way around. so, that's the way that they did that shot is they had a ring of like hundreds of cameras. So, then like he fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that that famous shot in in Matrix. But now with Atlas we can do this with just as few as through three cameras. So, like no studio capture, no green screen, no expensive calibration.
2:25 We can literally stick like three cameras on three iPhones on tripods, use these to take sort of a video of something happening, like someone shooting a basket, someone dropping a strawberry into a bowl of milk. And then from those three like iPhone videos, we can then reframe the shot, and imagine like a like freeze time, have the camera fly in like as the milk is splashing up, and get these amazing frozen time views. and we can do this with just just a couple cameras.
2:49 >> Can Can you just maybe What is the simplest description of what Atlas does? Like what goes in and what comes out? >> Yeah, so one of the one of the really core principles of Atlas, like the most fundamental thing, is it does new view prediction. and this is a a really fundamental primitive that we think is super exciting, a super new primitive for for base models that no one's ever done before. Right? So, we know LLMs are built on next token prediction, we've seen video models as being built on next frame prediction. Atlas is really new view prediction. Right? That given some number of views of a scene or a description of a scene, those go into what we call a spatial context that describes implicitly what is the world that we want to talk about, then you can point a virtual camera at at an arbitrary point in space and time, and Atlas will understand what that world is supposed to look like from that position in space and time.
3:34 >> You know, you know, Ben, that you know, with with a with a bajillion video models out there all claiming to be world models and all claiming to have novel views, and can you maybe tease apart can more concretely how this is different from like the myriad models that have come before? >> Mhm. Yeah, I think what Justin was saying about the spatial context aspect is super important here. So there's many video models, a lot of video models actually got their claim to fame from their single dimension input or their start to last frame interpolation. Now we're starting to see models that can do this kind of omni referencing with, you know, 20, 30, 50 images. But what's key with Atlas is that it actually has a kind of like specially grounded meaning to every frame you put into it. So it's not just an image that the model's going to interpret whatever way it wants or you can kind of try to argue with it in the the text prompting and get it to do something specific. With Atlas, every image actually has an associated three-dimensional camera pose and that means that you can perform this task of reconstruction with an extremely high degree of precision, right? So if we had four views of this room, one at each corner, you can put those into the model and then get an exact replication of everything you see in this room and it's not going to guess what's in the other corner, like the relationship between things. It's just going to reproduce exactly what you give it. And you can also do that in a kind of creative or imaginative sense, too. If you take two photos from different, you know, AI generations or real-world locations, you can actually position and stage those to build these kind of intentionally directed fly-throughs that are really governed by exactly the precise place that you put the content you want and where the camera's going to look and travel, which is very different, I think, than the kind of like more slot machine effect you get of having to retry generations over and over with just that kind of higher level of text control you get with video models.
5:15 >> Is this just kind of an obvious, you know, scaled-up version of a traditional video model or is it a new architecture? >> I think it's it's a pretty new thing for a couple of different reasons. One that we talk about is it does both generation and reconstruction jointly in the same model. Like Ben was saying, this thing can take a couple of views of this room and then reconstruct everything in this room exactly as you see it. And historically, reconstruction has been its own subfield in computer vision with its own specialized task, its own specialized models. And generation is what all the text all the text to video models are really good at like what all the all the big diffusion models we've seen the last couple years.
5:51 And those are great for creative applications. I want to imagine something that's never been there before. but now with Atlas for the first time we're putting these two different parts of visual intelligence together in one model. So it can do both 3D reconstruction and generation together in one architecture. So to do that we had to make a couple changes. one is we had to make it multimodal from the start. So this thing natively works on text, it works on images, it works on videos. It also works on camera poses as a native input to the model which I don't think anyone's ever done at the pre-training phase before. and it uses it uses 3D as a native modality that it works on. So this thing from the beginning was designed to be natively multimodal in a way that no one else I think >> Sorry, I just I don't know this space super well but 3D is this like depth or models and like what is that what is that mean?
6:38 >> Yeah, so the formulation we used so far is depth maps. >> Okay. >> Right? So right now you can have a when you have a frame that has a a a virtual camera telling its position in 3D space, yeah, that camera position and and camera parameters are a native input to the model. And then what attached to that camera position you can have both like RGB telling you what is that what is that position in space look like and you can have a depth map that tells you what is the spatial structure of that position in 3D space. So then you know text, image, video, 3D cameras are these modalities that this thing all does jointly in a multimodal way.
7:10 >> I want to add something cuz I think what Justin just said is actually so important and also what Ben said that it's under appreciated. It's the first time we have a unification of pixel generation and pixel reconstruction. In the world of computer vision this field has been around for more than half a century. sitting here having been in this field for decades, I cannot tell you how many PhD thesis have been written on the problem of reconstruction or novel view things synthesis and also our field traditionally have multiple tracks. You go to a computer vision conference, you have the pixel generation track, you have some recognition track and you have 3D reconstruction track. This is a elegant model that combines or unifies the problem of reconstruction and generation by anchoring on viewpoints and the the viewpoint estimation and that's just incredibly powerful.
8:15 >> Can Can you Can you maybe Well, can we take a step back and then maybe you just fill something out? So, when when you started the company, I remember you saying you know, you want to tackle spatial intelligence, right? And you know, now we have this new model. And so, it feel I mean, like as a layperson, it feels very general to me. You've got next product prediction and this is next new view prediction. >> New view prediction, right?
8:37 >> So, like you can get one view out of a set of views and you have a new view. Can maybe you pencil out like like how this is a significant step to this general problem of spatial intelligence? And maybe by like starting to describe what spatial intelligence is? >> Well, spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it. >> Yeah.
9:03 >> Now, we can argue is it 3D or 4D? Ultimately, it's 4D with the time dimension, but even just 3D, these are the fundamental tasks that one has to do or spatial intelligence has to enable. And then we talk about with that, you can render, you can simulate, and you can plan actions. But to do that, a fundamental problem to solve is to understand the geometry and structure and the physics of the space. >> Yeah. >> And I do believe Atlas is a significant step forward because now with every single frame, you have a you can generate an estimate a important piece of information, which is the the view viewpoint, the camera pose. And that is the most critical information one needs about the geometry of the of the space. And that can lead to all the emergent behaviors we see in our in the downstream of the model, which we showed in the blog. So So in the on the path to spatial intelligence, generating pixels is definitely a a early step, which we have seen with what you call it gazillions of models. But generating pixels that are truly spatially contextualized and grounded is absolutely another major step. And that is the very hard step that Atlas has taken.
10:27 I we definitely have, you know, we can just keep going here, right? Like there is the fourth dimension of time, which will bring in dynamics. And there is more higher fidelity simulation and delineation of the space. So this is part of the the road map of spatial intelligence. >> Great, yeah. I mean, I definitely want to like dig into like where this is going. But first, maybe let's talk about getting here. How long has World Models been in existence?
10:58 >> Two and a half Two and a half, yeah. >> And so you you've actually released models before. So what why didn't you just jump right to Atlas? >> >> Good question. >> It's so magic, right? Yeah. >> Justin's team needs a lot of chips. >> Yeah, we need a lot of GPUs to actually scale this thing up. So what one of the last year we released our Marble World model, and that was the first kind of big major World model that we put out. That that powers our current Marble product. And Marble and Marble is really cool. Marble can take images, it can take videos, it can take text prompts, and use these to generate 3D worlds. But one of the biggest differences between Marble and Atlas is exactly what is that output modality.
11:33 So, marble was really focused on Gaussian splats as an output representation. So, whatever you're inputting, it's going to output a 3D world represented as a as a Gaussian splat. And Gaussian splats are really useful, right? They're really nice, they're easy to render, they are they they can they can render efficiently on mobile devices, on VR devices, they can interoperate with other with a with game engines, with simulation engines. There's a lot of nice things about Gaussian splats. But, you know, I think that was kind of a bottleneck in the previous marble model. So, what we did with Atlas is redesign the thing bit.
12:01 and we realized that we need to bifurcate these modalities earlier and actually have these things all these modalities working in a more unified way in the model. So, now with Atlas, the the the fundamental primitive is not like a good generate a Gaussian splat world. The fundamental primitive is is as we said new view prediction. and that can generate RGB frames, that can generate 3D, and we can use those to generate a beautiful Gaussian splat worlds when you need them. but, we don't need to bottleneck our our our outputs through the Gaussian splats when we don't need to. And that was a that actually took a lot of, you know, blood, sweat, and tears to understand like what are all the pros and cons of these different representations. So, that's one part of it. the other part is you got to like climb the scaling ladder, right? You got to like work your way up and like do smaller experiments, do smaller models like to build your conviction on what's going to work and what's going to scale. and there's there's you know, if you could instantly know the right thing that's going to scale, you know, you should just do that. But, when we when we started the company, the world was a very different place. Like there's there's no scaling law of special intelligence. Right. So, like when we started the company, like the world was in a very different place, the tech was in a very different place. We had a lot of ambitions for where we wanted it to go, but it took a couple it took a a couple iterations for us to hit upon this formulation that we thought is actually is actually like this is the one. This is this is the one that can scale out.
13:16 >> You know, Ben, you know, being the creator of Nerf and doing a lot of 3D and reconstruction, so it's not so obvious to me that like if you have multiple views that you actually end up with a 3D thing. But, like you've kind of like made a career of ending up with a 3D thing. So, maybe talk a little bit about like kind of that step. >> Yeah. Yeah. yeah, I mean as you said, I've spent many, many years of my career as a mass majority of my career actually working on producing 3D things from images.
13:41 and this is actually something we we talked about a lot early on in the company even of like is this going to be the approach that produces 3D, right? Are we going to synthesize multiple views and then build 3D out of that? Are we going to try to go direct to 3D? Like there's been a lot of uncertainty in the field around like which of those approaches kind of will win out or will kind of like reap the best advantages earlier on. but I did have a lot of conviction just from seeing the kind of power of what I would almost call the brute force scaling scaling at a very, very, very small baby scale, not like real model scaling, but the scaling of dense reconstruction that we had seen happening over the past 3 years before.
14:17 so, basically we put up >> What Why is dense reconstruction Yeah. dense? >> Yeah, dense. >> Dense dense is >> Because I know we're going to talk about sparse and I want to make sure that people understand what is dense and what is sparse. >> Yeah, so I think this is actually even on the kind of like business and commercial side I I think been one of the challenges of productizing 3D reconstruction technology like at a fundamental level, right? People kind of don't in in a in a casual sense like you think I took three photos of this object or I took six photos of this room. Like I I look at the photos, I can understand in my mind like how this piece together. I can kind of fill in the gaps and get it, but there's just never been really any kind of reconciliation between those like really data-driven priors and then the kind of brute force dense reconstruction, which it actually is much more akin to almost like scientific or medical imaging what we did in dense reconstruction, right? You basically have to say every single thing I want to appear in this reconstruction, I need at least three or four views of it. And if you think about that, like even just in this room, right? There's like under the microphone, under the table, between every different crack and crevice and the plant leaves, right? To actually truly get a picture that covers every one of those spots, it's this like very tedious and exhaustive effort to walk around the room. I think you've all seen me running around various places like capturing them. It takes, you know, for for someone who's well-trained per se, like it can take minutes, but if you hand a casual consumer or even some like kind of professional trying to do this for the first time, an average like cell phone camera or capture device, like it's going to take them probably an hour. I've seen someone for the first time trying to scan a multi-room environment spend like two hours walking through it and get enough coverage. And that's just this like very, very exhaustive and tedious loop.
15:59 And so yeah, when we say dense, we really mean dense. It's like this room I want >> You just need like lots of >> so many photos, right? I want like 100, 200, 300 photos of this room to to capture it. And what we're trying to do is bring that down to like three. >> Three. >> Right? Well, we're saying like 50, 100 x reduction. And then that's at that scale where it just completely like flips that calculus on its head of like what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and like kind of unearth a lot of footage in the past you would never have treated as reconstructible and go back and like bring it to life as 3D potentially. This is something we've been playing around with a lot with Atlas, right? Like taking old clips. Like I've taken a bunch of my own old captures that never worked before and then put them through the system and I kind of seen a reconstruction for the first time or taken my old captures and thrown away 95% of the photos I took and you know, imagine angles that I never would have gotten from a traditional kind of like Nerfer Splat type reconstruction.
16:55 >> One thing that's under appreciated on the website of the demos is the Stanford demo where Ben showed anywhere between 3 to 25 images you can reconstruct that entire Stanford quad. But the thing is we had to show it from aerial view. But every single input image is been standing on the ground taking a picture from the ground. So, everything you see are generated but according to the laws of reconstruction.
17:27 And this is really magical. >> And this is where like generation and reconstruction need to interplay in a really fundamental way to solve this problem. Because under the classic kind of reconstruction stuff that Ben was talking about, like the reason you need so many views is because I need like multiple images and I need to triangulate this point in 3D space and see it from multiple viewpoints. So, that means like that's required in the traditional version. And on the flip side, anything that wasn't captured in these views, like any pixel that was not visible in one of the input views will be a hole in a 3D reconstruction.
17:59 Because fundamentally like if a thing wasn't visible in the input views, you know, you need to imagine it to fill in the gaps. And that's fundamentally a generative process. So, even in this room, even if we set Ben loose with a DSLR and like let him like capture like hundreds of views of this room, even the world expert on doing these dense captures is still going to miss some spots. Like he's not going to get like underneath all of the microphones or underneath all the tables or like in between all the chair legs, you're always going to miss something, no matter how many views you get. So, that that's where you need generation as another mechanism in the model. Because you're never going to get everything.
18:31 So, you need to have some generative capacity for the model to imagine, oh, based on what I'm seeing, then like first triangulate what I what I can see, but then fill in the gaps of the stuff that inevitably inevitably was not captured. >> Yeah, and there's something like super cool about this that LLMs have really understood this for a for a long time, right? There was kind of one of these like context wars like the first couple years I was like, oh, we got to 128 to 256 to like 512. We got a million, right? And everyone kind of understands now at a pretty tangible level the value of, you know, you crank your context length to high when you're using your coding models. It's a hard problem. Like everyone has a feel for that. But like no one has pushed that at all on the image and video model side in the same kind of like principled way. Like no one's out there trying to like put a like an hour-long video through and do a needle in a haystack retrieval of like a frame at the 30 37-minute mark. Whereas with reconstruction and generation, you actually have the same exact thing of like reconstruction is just like generation with a really long context and you put a lot of stuff in it. Right?
19:24 Like that's the way to actually build this continuum where you kind of bridge between those two things. And like Atlas, like being able to like this is something we could never do with Marvel. Marvel had this kind of fundamental blocker of like you couldn't really jam more than honestly like a couple images in. But Atlas, I can go and I can actually take like a 64-image capture and do like a fly-through of an entire house and everything is grounded by being, you know, seen or like almost seen or like slightly extrapolated from what's not there. But you're just getting these, you know, I'm taking captures I did with 2,000 images of a multi-room house and taking it down to like 30, 40 inputs and the fly-through looks like basically the same. And this is just like totally inconceivable before and it's all enabled by building this gracefully scaling kind of context window that you can dump stuff into.
20:06 >> And so the way to think about it is like the the sparseness are the pictures that you physically took and then Atlas is a model creates the rest of the views and then you use classic reconstruction techniques. Is that roughly the way to think about it or >> In some sense, yeah, yeah. I mean, that's the beauty of Atlas is like you can take however many inputs you have down to like a single view and then you can always use Atlas as this this rendering engine to produce anything else you want. Right? You can you can navigate it like a virtual camera. Yeah, exactly. You can just you can say like, "Okay, I have a picture here. I want a picture there, there, there." You can make a couple of those then you can say to dense fly-through. You can do this in sequence because it's an auto regressive model. It's up to you, right, to kind of pick and choose what you add interactively into the context as you generate.
20:49 >> I mean, the thing that I just blows my mind is Listen, I just have a very simple mental mental model. I I I I four pictures and then I've got to like have the model extrapolate between them and then it has to fit when you reconstruct. Like it's got to be 3D. Like and I always think of these diffusion models as like being visually great but not accurate. And so like And and I don't even know if there's a question here, but like how how how is I like the room fits? So like how is it that it's three 3D consistent? Is it just lots of data or >> Yeah, I mean it's a part partially it's a belief in the scaling hypothesis, right? Like, you know, >> Did you By the way, but I have to ask, when you set out to do this, did you know it was going to work?
21:32 >> I was pretty sure. >> >> Were you sure? >> I think three of us have total conviction about the scaling law. That that I think we do. I do think the exact architecture choices and data mixtures is where the the devils are in the details. I, you know, have watched Justin and his team going from we really don't know how long this is going to take to oh, maybe sign of life to wow, this is going to work. So it it no one what no one has done it, but I think the hypothesis, two hypothesis, one is scaling law hypothesis, the other one is next viewpoint prediction. We had conviction of these two things primarily.
22:20 >> So I think I was very convicted that it was going to work. I was not sure it was going to work at this well at this fast, right? Like I thought there's a chance that we do this. Maybe it's not clear that like the first cycle of pre-training a new model with a new architecture and a new paradigm. Like the first cycle of that working is insane. So I thought there was a chance in which we had to we might have had to do a couple more turns of that of that pre-training cycle before we got to the level of quality we we wanted.
22:44 >> Is there is it are we kind of like at the end of like the scaling for this architecture approach? We need another breakthrough or is there >> No, no, we're at the beginning. >> Really? >> Yeah. >> Without without changing the architecture? >> Yeah, we're basically at the beginning. I think we're basically at the beginning and we're basically limited by compute at this point. All right, like data is is important as Feifei likes to point out, but like everything has a bottleneck and I think the main bottleneck on continuing to scale this thing is actually training compute.
23:07 Right? Like during development, we trained a sequence of models. We wrote about this in the blog post a little bit, but we trained a couple of models that like the first couple of runs of the scaling ladder. and each time we made the model bigger and each time we trained it for longer, each time we put it on more chips, like it got significantly better. And the model size that we like the model that we showed in the blog post is obviously the biggest and best one that we trained, but the thing that was limiting it was not the scale or the data or anything like that.
23:32 It was literally like we had a deadline of when we wanted to release this thing and therefore we backed up what we could afford to train in time for that deadline. >> But here's a little bit of a insider story, right? Like Justin and team are are training from the smaller and slightly bigger, you know, are train having these roadmaps. And then there was one day in summer, early summer, that it's not even this the current Atlas model size, it's a smaller model. And then Ben, Justin, Ben feeded into, you know, the viewpoint generation. And remember that famous table, the garden table for Nerf paper and many papers, that overnight I got a slack. I mean, we all saw the slack from Ben that our camera flew through under the table.
24:19 >> With the soccer ball. >> yes, with the soccer ball. >> ball emergent or is that in the original picture? >> It's It's It was real, right? >> Okay. >> That morning the three of us looked at each other in the eyes and say, "That's it. This is We're going to build this." Like we we made a decision within 5 seconds. How This is just a no one has ever seen this result. >> Ben, can you talk through maybe more specifically the use cases? So so World Labs has historically had a lot of users that were creatives and they use it for like consistency in, you know, like whatever 2D images and for movies and for 3D and for games etc. And so maybe can you talk about how this extends use cases or cater to the existing ones and then we'll I'd like to talk about robotics actually.
25:04 >> Yeah, sure. yeah, I mean it's kind of funny actually one of the kind of main ways we even saw people using Marvel plays exactly into this new view prediction case. Like a lot of our >> Marvel being the previous sorry our Marvel our previous product. >> like people would take that product, put an image in, get a full 3D scene as a Gaussian splat, take a couple of screenshots of it from different points of view and leave. Right? And we're like We can just make these images and that's generative AI cool, right? So So I think like and you know, there's a lot of degradation there. They're like, "Oh, this splat could look better." And it's like, "Okay, what if we just generatively model those viewpoints with that exact modality of control?" So I think like even that core capability of just like view synthesis, it's sort of been this academic problem for a long time. But in the sense of "Oh, you're going to do this really dense capture."
25:48 Like like generative view synthesis is a relatively quite a new problem. And we just see so many people who in this creative pipeline, right? People have a multi-stage workflow, right? I don't think there's a single person out there using one monolithic model, not even C dance or whatever for for their entire task. people will have this like, you know, kind of a bunch of storyboards and mood boards of images they pull out from like their favorite collection of image models. And then they'll go to different video tools and like build those together as keyframes. Then they'll go and like clip and edit those later, right? So we were seeing this like sort of, you know, niche but very specific use case for Marvel as just providing that like sanity that you can ground your generations in some kind of 3D consistent world, right? People, you know, I don't I I fought with with various image models to ask them to like give me different viewpoints of a room.
26:33 And every time you can just look and see, "Oh, things kind of moved around. Like it's not stable." And like even that one seed of a use case I think kind of signals that there's this value and there's hiding under the surface there. Like there's just decades of people being used to persistent 3D state like virtually modeling what they would be doing in the real world and having, you know, a stage and props and like elements there, whether it is for a movie or a show or a marketing shot or like building out game environments.
27:00 Like this this statefulness and persistence is so key in how people think about spatial reasoning and like developing an environment over time. Like people don't think in this ephemeral like generate a thing, generate a thing, like just throw it away, keep my text prompts. Like people want to build this like collection of assets and like model a world in that way. So we're we're trying to provide like again with with the spatial context mechanism and other things like we're trying to provide that level of control and precision and the ability to adjust different modalities of input starting with the post images, but you know, we want to give people more control over the elements of the things in the scenes they're looking at and editing and interaction and all that as we go forward. And I think that that it unlocks like further use cases in those areas we're already seeing, but also expanding out into kind of any place people want to create a virtual replication or or like, you know, a pre-imagination of a real-world space they need to build, right? For architecture and construction. Like I talked to a guy at some point building booths for conferences, right? There's just so many things in the world you don't think about need to be fabricated and every single one of those basically goes through this like pretty painstaking virtual design phase. And of that process like the part where you go into 3D software is kind of one of the most like arduous and like labor-intensive parts right now. Like like taking feedback on a 3D design from kind of like verbal commentary or sketch or really really quick stuff you got from like a creative director or like a design director or an architect or whatever. Like mapping that back into the 3D representation is like 95% of the work, right? You can have a meeting get feedback and then you go back and do a week of revisions. And that's just because like our software is kind of decades old at this point and it's it's just never became as intuitive as, you know, playing with Legos or like pottery or doing this stuff with your hands or sketching with a pencil. And this is the real place where AI can actually unlock a ton of value for people in their process, whether it's a creative application or something more industrial or design or whatever.
28:55 and that like really motivates me to kind of build different flavors of of our model to cater to those kind of people. >> Yeah, I I I can understand how it helps with the creatives cuz like Marvel did that. And also how that extends to things like designer architecture. Baidu you acquired a robotics company. And so it's >> >> it's less less >> talked about it too. We just talked about >> know. But it's the last time we talked to me especially in the context of Atlas like how that maps to robotics. So if you wouldn't mind just penciling that out.
29:22 >> Yeah, actually Atlas is a a key part of the puzzle. So, we acquired this company that was formerly known as Synnex. And what is their key technology? Right now their key technology is a system that goes from real to sim and and then sim to real. And what does that mean in robotic situation? You want to train a robotic arm to, you know, figure out how to, do cabling, let's say, in a in a industrial setting. Well, you need a whole bunch of data to first train a robotic policy to do these cable cables cabling activity. And then you want to evaluate if the robotic policy is doing a good job. And then you deploy the robot into the cabling environment. In order to train, what you what this company, Synnex and now our robotics team used to be doing is doing exactly what Baidu was saying, dense reconstruction. You take >> All right, so Synnex >> pictures of a situation and then and try to reconstruct that environment. It's excruciatingly painful. Takes a long time, laborious, and it really blocks the velocity of robotic simulation real to sim, right? So Atlas really is the next generation technology for that. And this is not just for robotics cabling or anything. We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data. Because it's so hard to collect real-world data where robots, you know, are operating and and in order to not only you need to collect the data of, let's say, the cabling situation or dishwashing situation, whatever. There is also a very important step called randomization. Is that you have to take the same environment and then randomize the conditions. So, the cable doesn't literally only, you know, bend this way. It can bend a different way or the box can have different sizes, colors, different lids, and all that. And and or in different parts of the scene. So, you have to go through a real-to-sim situation in order to get enough of that data in addition to other data you can get from internet. So, this real-to-sim step will be you know, really helped by by Atlas.
31:57 That's just the first part of this is meeting the the the robotics needs in the current technology cuz we don't yet have a a frontier foundation model that's robust enough for for robotics. But Atlas is a omni model. It's a multi multi-modal model. It takes on different kinds of input and generates different kind of output. You can totally imagine the next step is Atlas taking in data that's in the in that's dynamical.
32:31 And that can really start to bridge the gap between, you know, action planning and and a robotics and and the Atlas output. So, that's all narrow map. What are you going to say something? >> Yeah, I I was going to say there's something fundamentally different about training a robotics policy compared to really any other application in AI we've seen before. and that's like if you're generating a piece of code, like you're generating an image, you're generating a video, I'm the model is fundamentally creating this this artifact. And that artifact like there's a lot of examples of artifacts that you can go out on the web or somewhere and collect. Right? You want to generate images, there's a lot of images out there. You want to generate videos, there's a lot of videos out there. You want to generate a code base, there's a lot of code bases out there you can learn from.
33:14 >> Yeah. >> A robotics policy is something fundamentally different. It's not producing a static thing. It's instead a policy that's going to go out into the world, make actions, and like try to achieve a goal. And the world's not always going to respond the way you expect, right? Unexpected stuff is going to happen. So, a robotics policy is like fundamentally an agent that is out in the real world interacting with the real world, and stuff happens. So, you need like a critical part of that is the those those policies during training need to be exposed to every possible thing that could go wrong during a deployment. and that's where simulation is really key for robotics, right? So, there's the then there's two angles on that. Like one is the kind of classical simulation. You can go out You can go and like go to your favorite physics engine and like try to imagine creatively as a human designer, what are all the scenarios that might happen in this in this in this when achieving this task, and then try to write explicit code that models them all. That That's one angle. And that's an interesting angle with coding agents. Like that actually gets supercharged, too. But there's another angle, which is try to more data-driven simulation. Right? Like maybe we can Can we have a learned model that can understand how the environment how the world is going to respond to actions? And maybe it might respond in unexpected ways sometimes. Then could we build these neural simulators that are trained on as much data as we can, then use these neural simulators, these learned neural simulators, you know, as a simulation bed to train robotic policies.
34:37 and that's that's a really interesting future direction of Atlas. But but then it doesn't stop there, right? So >> >> so but once you have, you know, this learned simulator, like this learned simulator kind of already has in its like mental brain, like it understands the world, it understands how the world is going to respond to actions, and why doesn't the simulator itself become the planner? Right? Like the same That's that's kind of the core thesis that we've had around world models and their generality, that there's some core stuff that a model should understand around generating worlds, simulating them, understanding how they appear in different situations, and, you know, understanding how the world's going to respond to an action is highly related to imagining what kind of action I need to take to make the world respond in a particular way.
35:21 >> Yeah. What what one piece of feedback that I got I've been By the way, congrats on the launch. It was overwhelmingly positive. I think it was probably the most significant model launch this year. And yeah, everybody said glowing things. But one person who's an expert in the space who I texted was like, "What do you think?" And great person says, "It's fantastic, it's amazing, but there needs to be more dynamics." >> >> And so it would seem to be at least in the robotics case, but generally it was kind of ideal to actually have a world that moves. And so maybe talk a little bit about that and then any other future directions that A, you're comfortable sharing, but you think are worth talking through.
35:56 >> Yeah, I mean, like dynamics is clearly going to happen. Like actually >> We have baby dynamics. >> We actually do have baby dynamics already. And this is something I think people didn't quite appreciate, we didn't really highlight in the blog post, but like the previous marble world model, it was like fundamentally static. Like the the model just like could not handle any dynamics at all. And that was just like baked into the model architecture, baked into the training, like the whole thing was fundamentally static.
36:17 >> Yeah. >> we already knew that that was a big problem post marble, and we already fixed it in Atlas, right? Like the Atlas architecture is already fundamentally supports dynamics. the Atlas training data fundamentally has dynamics. And if you look carefully in some of the videos that we've even posted, >> It actually is. >> I saw I saw you see that. >> Yeah, the waves, the water waves. >> Yeah, so like some of the examples there's like waves in the water, like in some of the like air generated aerial views, there's like little cars moving around. So like dynamics is actually already in this model.
36:45 >> But is it By the way, dynamics seems very problematic to me if you're trying to reconstruct 3D from multiple views, right? And so like are these things like at odds or >> it's actually one of our thesis here is that like, you know, if you're just going to do fundamental 3D reconstruction, you actually want to have no dynamics. Like you want to be able to model like exact views of the scene with exact frozen time.
37:03 >> Right. >> but then like this is actually kind of a problem with our previous Marble approach, right? Like there like you can try to find data that's fully static, but that's really hard to scale and really hard to get more of. And the thing we realized is that even in the case where I want a static output in the end, the best way to get it is actually expose the model to dynamics, right? Like expose the model to as much dynamic stuff as you got, as much static stuff as you got, and let the model figure out how to factor out the dynamic stuff. So in especially in like the This is like >> And this is actually So again, in the Atlas pre-training already, like it it saw a ton of dynamics in the pre-training already. Then the post-training that we did specific to this checkpoint in this release was focused a lot more on on static stuff, focused a lot more on spatial movement and less on temporal. But like we already have we already like I'm pretty sure this the the the pre-trained checkpoint already has a lot of latent dynamics in it.
37:52 >> Yeah. >> and this is something we're going to improve quite a lot going forward. >> So so Ben, does that mean we're going to get 4D video? You can like go walk around. >> I mean, I think >> see the smile on their face. >> So I I I actually can see if like you just stopped now and you only did kind of bigger, you know, faster, better, you could build almost an entire industry. Like I feel it feels like a very horizontal primitive. And then and if you did nothing else, but are there other things that are not just kind of bigger, faster that you're excited about with the applications you're focused on, which tend to be kind of more on the kind of content creator 3D side.
38:24 >> Yeah, I'm really excited about pushing that kind of multimodal aspect. I think different modes of control is so critical here. I think like it's super under appreciated, especially in the academic community, how critical it is to add control conditioning to these models to kind of get out what's inside. I mean, honestly, this >> I'm not trying to I don't even understand what those words >> So, translate in layman's language is so editability. I think editability is is the key here.
38:48 >> Yeah, so I mean, this is something we've seen in like sort of single image models and starting this year in video models is starting to be unlocked in terms of oh, like getting that flavor of like multi-turn or really like intuitively interpreting like I want like this person and this object and this thing to happen and kind of combining those all together in like one pastiche without having to do a lot of like manual work with the system. Like it just interprets it like kind of frontier image models are kind of there, right? For in terms of editing. But, we haven't seen that propagate out as strongly into video yet and then into world models, right? We've seen some really kind of toy examples of oh, I can like put in a sentence and like, you know, a dinosaur appears or something with these like sort of real-time models. but, I want to like turn that up to really industrial strength and make that cuz like the the trick here is you got to add control, but not compromise the quality of the model or it just becomes a a party trick, basically. Like it's like no one is going to seriously think about swapping their like cutting-edge frontier video model usage for your model if you give them extra knobs, but the quality degrades. So, I think it's really that game of like how can we maintain maintain like the high bar we've set with the outputs we're able to get in the current model and then add all kinds of interesting stuff that people will ask us for in terms of like I want to interact with the scene or control the layout or control like the identity of the objects and the things that we're seeing within there or control time, right? and I think that's like an axis where it opens up like a ton of really interesting product and interface work. The more complexity you add there and richness in terms of kind of like enabling you to really think about like redesigning almost from scratch the way people interact with sort of like stateful, you know, 3D worlds in the computer. Like that's that's really the end goal here. That is getting like all the capabilities you need to build that kind of system.
40:28 >> Awesome. Anything to get out of that as far as new functionality that you'd be excited about that's not just bigger better? >> I think for me, let's go back to the first principle of intelligence. Intelligence is not sitting there stuck and just seeing something or interpreting something when it comes to space and physical space, right? It's really this closing the loop between seeing and experiencing and interaction. So, just thinking about going up that ladder is exactly what Ben said.
40:56 >> Cool. I think one interesting notion there is this notion of AI completeness. >> You heard of that before? >> Yeah, yeah, I have. Yeah. >> So, like everyone >> Yeah, I I hear about AI complete, by the way, in terms of LLMs, which is like you have to be basically, you know, like the smartest LLM to answer the question what the smartest LLM will need to answer or you have to solve general intelligence. >> No, no, it's it's basically it's it's a connection to Turing completeness, right? Like the idea being that like a task is Turing complete like in classical complexity theory if like I can take any class of any problem in this category, reduce to that one problem, right? Three SAT is a classic example, right? So, you can take any NP-hard problem and reduce it to three SAT. Therefore, therefore, you can use three SAT to solve any any problem.
41:33 >> Yeah, it's it's a yeah, yeah. >> So, so then like the the kind of like soft definition of AI completeness is like there's this fundamental primitive that's that's an AI task, but if I could solve this AI task in its full broadest generality >> You solve that. >> it would solve any intelligence problem. And like the classic example of LLMs is like next token prediction is AI complete because I could like, you know, there's the classic example, I think from Ilya, where like I there's a mystery novel and like the thing has to read the whole mystery novel and the final sentence of the mystery novel is like, "And the killer was Predict the next token.
42:02 So, like you could basically like frame any kind of intelligence task in terms of that. So, clearly next token prediction is something that people believe is AI complete. >> Yeah, yeah. >> But I think that something we're kind of realizing and Ben was talking about this earlier today is like new view prediction, this primitive that we have in Atlas, especially generative new new view new view prediction. This is also AI complete. Right? And because I could take something like >> You could you could have the the movie and you do all of the frames of the movie and then like the killer walks out and then you predict exactly who walks out.
42:30 >> >> Exactly. Not just that, but I could say like I want to have a world where like Martine is like writing a proof of the Riemann hypothesis on the >> So okay, so to take a evolutionary view, right? That new viewpoint prediction is exactly evolution had to solve by making animals move. You you nature give animals eyes. But nature didn't give trees eye.
43:01 Eyes. Why? Because when you move, you see a new viewpoint. And that is the the the whether you call it AI complete or intelligence complete. So So we do believe very strongly that next viewpoint prediction is is the equivalent of next token prediction. >> Amazing. Well, with that, congratulations all of you on a phenomenal model launch. We're very excited for future model launches and thanks for coming. >> Thank you.
Summary
- Atlas introduces new view prediction, allowing for the generation of video frames and 3D reconstructions from limited camera inputs.
- The model can create complex scenes with just three cameras, eliminating the need for extensive studio setups or calibration.
- It combines generation and reconstruction capabilities, marking a significant shift in how spatial intelligence is approached.
- Atlas supports multimodal inputs, including images, videos, and camera poses, enhancing its versatility.
- The model is designed to handle dynamic elements, paving the way for future advancements in 4D video and interactive environments.
- Atlas aims to streamline workflows in creative industries, architecture, and robotics by providing a more intuitive and efficient way to create and manipulate 3D spaces.
- The technology has potential applications in robotics, enabling better training and simulation of robotic policies through realistic environment reconstructions.
- Future developments will focus on enhancing dynamics and interactivity, allowing users to engage with and edit 3D environments more intuitively.
Questions Answered
What is Atlas and how does it contribute to spatial intelligence?
Atlas represents a significant advancement in spatial intelligence by enabling the generation of spatially contextualized pixels. Unlike traditional models that rely on extensive camera setups, Atlas can achieve similar results with minimal equipment, thus unlocking new possibilities in video generation.
What is spatial intelligence and why is it important?
Spatial intelligence involves generating, reasoning about, and interacting with space, which includes understanding geometry, structure, and physics. Atlas makes strides in this area by providing crucial information about camera viewpoints, essential for simulating and rendering environments.
How do generation and reconstruction work together in spatial models?
Generation and reconstruction are interdependent processes in spatial modeling. While reconstruction requires multiple viewpoints to triangulate 3D points, generation fills in gaps where data is missing, ensuring a complete representation of the environment.
Why is a persistent 3D state important in creative applications?
A persistent 3D state allows creators to build and manipulate environments over time, which is crucial for industries like film and gaming. Atlas aims to provide stability and consistency in 3D modeling, enabling users to create coherent and evolving spatial narratives.
What are the future directions for Atlas and the importance of dynamics?
Future developments for Atlas will focus on incorporating more dynamic elements into spatial models. While current capabilities include basic dynamics, enhancing this aspect is crucial for applications in robotics and other fields where understanding movement is essential.