Section Insights
Introduction to Cooking with Codeex
What is the Cooking with Codeex workshop about?
The workshop introduces attendees to the capabilities of Codeex, showcasing various projects and applications that participants have created using the tool.
- Codeex is being used for diverse projects, including synth plugins and video editing.
- Gaming is highlighted as a strong area for Codeex applications.
- Innovative projects like self-driving golf carts demonstrate the versatility of Codeex.
Live Demo of Codeex Capabilities
How can Codeex assist in building applications?
The speaker shares their experience using Codeex to build a Mac OS app, emphasizing its ability to provide best practices and context, even for those unfamiliar with certain technologies.
- Codeex can help users unfamiliar with specific programming tasks by providing guidance.
- The memory feature allows Codeex to learn user preferences over time.
- Live demos can be unpredictable, showcasing the real-time challenges of app development.
Deep Integration with Browsers
What advantages does Codeex offer for web application development?
Codeex provides deep integration with browser engines, making it particularly effective for testing local web applications and interacting with software that lacks a programmatic interface.
- Codeex's in-app browser is beneficial for testing interaction designs.
- It excels in automating tasks in legacy software without exposed APIs.
- The Chrome extension enhances the ability to manage authentication and state.
Collaborative Development with Codeex
How can Codeex facilitate collaboration in project development?
The speaker describes using Codeex to manage permissions and collaborate on projects, including creating a customized form builder for internal use.
- Codeex can help automate the management of deployment keys and permissions.
- It allows for the customization of existing tools to better fit team needs.
- Collaboration features enable sharing and managing responses within a team.
Automating Feedback and Development Processes
How can Codeex streamline feedback and development workflows?
The speaker illustrates how Codeex can automate the handling of feedback from various sources, categorizing it and initiating appropriate actions based on the type of feedback received.
- Codeex can classify feedback into categories like compliments, bug reports, and feature requests.
- It can automate the creation of threads and pull requests based on feedback.
- The integration of multiple Codeex threads can enhance team collaboration and efficiency.
Transcript
0:18 All right. >> Good morning, everybody. >> thanks for bearing with us. There was a long line outside, so we just wanted to make sure everybody could get in and get a seat. thank you so much for being here. This is the cooking with Codeex workshop. We're so excited to have you. my name is Charlie. I'm on the developer experience team at OpenAI. >> And my name is Gabriel, also on the developer experience team, flying in from Singapore. Nice to be here.
0:48 so before we kind of dive into the nuts and bolts, you know, I wanted to do a quick review of what have folks been cooking with Codeex. >> It is a cooking show. >> these you might have seen some of these tweets. but you know there's a real range of stuff that people have been making. Some of it pretty pretty incredible. I think this one a session with nearly 300 sub aents. Has anybody beat that in a given session? No. I think I think Dom might have, but we'll see.
1:27 Codeex working on a synth plugin for Ableton. really really cool stuff. I think on our team we have a team member Brent who's doing at this point video editing with codeex driving premiere right which is not what you would normally expect to see from from a coding app. playable you know games I think like gaming to me is is really a place where where the models and the app can shine.
1:59 the ability to go and check your work for the game as it's done or to hand off between different threads and you know manage audio versus art design versus game mechanics versus you know difficulty balancing. I think I think the different threads and sub aents are amazing here. And last but not least you know building a self-driving golf cart with codeex which I think is just absolutely incredible. it kind of makes me wish I was back at school so I could be hacking on stuff in my nights and weekends again.
2:38 But where have people been cooking? I think that's pretty fun. So you we sort of see it everywhere. Everyone's codex mixing everywhere, anywhere, all the time. I mean, literally in between meetings, you've seen people like holding their laptops everywhere, keeping it open. It's become a meme at this point. Fortunately, we now have Codex on mobile. So, this gentleman in Japan was able to like bounce between ideas in bouncing between cities while on the bullet train or while walking a dog.
3:08 And you don't just ha you don't necessarily have to use Codex on the Codex app or use it within like the chat GPD app. You can just put Codex anywhere with it being open source there being the app server which we'll talk about briefly at the end. You can put it here. So this Daniel for example put it into three e- in devices. So you really can codex max anywhere. So if it's your first time sort of using codeex do check it out. Do give it a download. What's really nice is that if you have existing configuration for other agents you can just bring it in and just get started really cook quick.
3:43 And today we want to through the cooking show help you cook really fast and cook something really cool. So there's a quick agenda. We're going to talk about what's new and then the real pun intended meat of the session is the six ways to cook with codeex. Some closing thoughts and then the second hour of the session we'll have like a mini hackathon where we're going to have live cooking right in this room and we'll be like walking around helping to be your sue chef to your main chef. So what's new?
4:11 What's fresh from the oven? A few a bunch of things I think fundamentally to just really help you give context to Codex really fast. There was record and replay. You spend time maybe describing, oh, you got to press that button, you got to press this button. Maybe just show codeexc by doing it. It builds on sort of our technology with computer use and so forth. Tread handoff really fun. You can sort of push things from the cloud to the remote to your local and vice versa. We'll demonstrate some of that. And the mobile app is now in generally available. So now there's like sort of onetoone device pairing, a much more secure way of working with codecs.
4:48 So, it's it's been 2026. This slide barely summarizes everything we've shipped and we're only 50% literally 50% of the year. Some things have been really fun. The plugins, computer use, multimodality, being able to use Codex on the mobile, app shots, that's a personal favorite. These things really make Codex and the Codex app really fun to use. And we'll be talking about some of these. They are all sort of key ingredients. So, that's cooking a Codex. we wanted to like start with this quote where it's from our favorite movie Ratatouille. Anyone can cook but only the fearless can be great. And today with the sort of ingredients we're going to share with you some of the recipes but with your ingenuity we think you're going to cook some really fun stuff.
5:34 And these are the six steps to cooking with CEX. Number one you got to give Codex your context. There's a lot of context in like your documents, in your brain, in your conversations, in your email, in notion, linear, so forth. You got to give codeex that context. That's a really important ingredient. And then once you have it, you can literally just build things. And that's already just a step two. For the power users in the room, there's some interesting stuff where Codex can actually take action.
6:01 You might notice right now like on I'm on Chrome and you just see this thing nice thing at the top saying Codex started debugging this browser. what's happening in background is actually codex is reading the slides as I'm speaking and then you can sort of collaborate with codex on longrunning tasks and you want to work over time where you want to put this make it repeatable sort of ensure consistency and really like build those loops and finally we want to bring codex everywhere if you want to embed it in your own product or in your own internal platforms that's also possible so let's just dive right in for the first two steps cooking a codeex where you give it your context and you can just build So, a few days ago, like on Friday, Charlie and Christine and myself, we were thinking, "Yikes, it's it's Friday.
6:47 It's Friday night." We sent it at 4 4:44 p.m. We don't have a demo idea for this workshop. we had any suggestions. Then Christine suggested, "Hey, let's build a real-time translator." Charlie suggested, "Hey, let's use GPT realtime translate. It's one of our latest models. It's really good." And then we can even have a plugin. I've never built a Mac OS app, but with the plug-in, all the best practices are there. So maybe you would like copy paste, you would like Google or search or what these things are. Then you go to codeex, but I'm going to do something fun. So I'm going to take my two thumbs over here, my my right thumb and my left thumb. I'm going to put it on the command key. I'm just going to press it.
7:22 And bam, that's a app. All that context on Slack, all the conversation, that brainstorming is just right there. And let's just ask Codex. Hey, can you just build this? Yeah, cooking show is over. No, I'm kidding. And we're going to let Codex cook on that. And what happens with a appshot is that when you take a it's like a screenshot, but all the text is embedded as well. And it's very nice because that's where you have a lot of personal context, the conversations, there's brainstorming, you don't have to like repeat yourself. You just give codecs and you can like just figure it out. So, we're going to look let let that work on it for a bit and we're going to try something else as well. So, we have the agents SDK. It's open source. It's pretty nice. And you got like some things that have come up in the recent months. So, I was thinking, hey, maybe we could build like a website to sort of just briefly describe what's new and we're going to use this thing called sites. So, sites is something that we launched earlier in the month in June for in preview for enterprise and business plans. It lets you sort of create a simple site and it's really about just creating artifacts you can share. So here we sort I worked with Codex for a bit and let's just see what Codex cooked here. So we're going to open it up and it's all what's new with the agents SDK. You have the it's pretty nice. You have like six changes, sandbox agents, codex as a tool, tool search, realtimes agents. We'll be talking about that in a in the next few days as well.
8:43 And then you can see like the code snippets are here. So yeah, it's pretty nice with sites. You can then create artifacts that you can share with your team. So that's another example of bringing context because you have a code base. It's a really large codebase. There's git commits, there's history to it. is a lot of things. That's that context and that's step one. You're just going to give it to Codex to build it and that's step two. You can just build things. So let's give in that last last example here. Maybe like now we're Charlie's out on the streets think taking a walk with his kids, his dog, and he's wondering, hey, maybe you wanted to build something. So I'll let Charlie take the stage here.
9:17 Got it. Cool. so what I want to show you is Codex remote. Gabe mentioned we might be out in the world, right? We might be, you know, on the beach. we might be camping something but we still want to get work done. We still want to keep cooking. So Codex remote is a feature that's now available in chatbt a mobile app and allows us to connect to a you know laptop or remote server and issue commands that can be executed remotely. So in this case let's say we want to visualize something about this agents SDK right for me personally you know last weekend I was building a bunch of IKEA furniture so that's kind of on the brain. why don't we try to visualize how this repo works in the style of an IKEA manual?
10:00 Can you create a visualization of this repository in the style of an IKEA instruction manual? I want to understand how all of the primitives fit together. So, this is going to send for my phone. And here we go. So, you can use remote to keep working. you know, you leave your laptop on. Codex will keep the the laptop open while it's cooking in the background. and it means that, you know, you can sort of free yourself from having to be at the computer all the time or having to walk around with your laptop open all the time, which is 100% a thing that I've been doing lately.
10:37 >> It's almost like a regular occurrence. Every weekend, everyone's like leaving their laptop in the room at home and just like walking out touching grass, you know, still checking in. Hey Codex, is everything okay? Is there anything I can do to help unblock you? And in a sense, everyone's become a manager. So, let's see. Go back to the live translation. We'll see what's going on over here. All right, it's building. It's did some web search. It's finding out what's going on. So, I think we've sort of covered very briefly to another concept called plugins. So, just going to the slides. So, we go there.
11:14 Here it is. Yeah. So plugins are a bundle that sort of bundle together skills which are best practices, external prompts, maybe practices that are custom to your team, apps which have the configuration and provide the connectivity to external services and MCP servers which provide the tools that are necessary and wow already sort of popped it up there. We'll just skip that in a bit. So that's plugins. And if you just take a look at some of the plugins we were using here.
11:46 So if you just go over to the Codex app and you just scroll right up here, there's this button called plugins. And you'll see a bunch of plugins. So for example, what we were using was the built Mac OS plugin which has a bunch of best practices in terms of the skills for how to use AppKit, how to use liquid glass, how to sign and inspect the built application. And we were also actually using the OpenAI developers plugin. So it's 2026. we're not going to go to the developer platform and click create API key, save it somewhere, pray that you actually saved it somewhere.
12:18 Now with Codex and the plugin, it can just directly create the API key for you as necessary and it really keeps things really smooth as you it doesn't really interrupt you in that flow. So those were the sort of two plugins we were using. But there are other plugins that may be relevant like say Gmail, Outlook, maybe use Microsoft Teams, use Slack and that provides all the relevant connectivity you need for your agent to work to bring context into the kitchen.
12:44 >> I think I think I would add to that, you know, the Codex app comes with the skills to build your own plugins. So in giving the app context, a lot of times I will work through a difficult problem or a workflow that you know I'm not sure how I wasn't sure before how to tell Codeex how to do it. And when I get to the end of that conversation, I will often say, "Hey, make this a skill or even better make this a plugin so that to next time I need to do this, you can just invoke the plugin and get the job done and we don't have to step through all of this work again." So this is what Codex built in say 4 minutes and 2 seconds. Let's see if it works. So it's a live translation. So let's let's maybe do like I I understand there's some German speakers in the room. So let's go let's do German and let's just make sure the audio is working as well. And let's see what Codex can do. So again, I've never used built a Mac OS app. I'm a data scientist by training. Not very not the best at like sort of working with like web RTC websockets. But with the Codex app and the plugins, I'm able to get that sort of best practices through the through that. So let's just do that. Make sure the input change to the phone, the output is here. And all right, fingers crossed. This is a live app. It's it's built. So this is like we don't know how this is going to taste. And fingers really fingers crossed.
13:57 Oops. Hello there. Hello. Mic test. One, two, three. Okay, let's change the audio.
14:28 You know it's a real demo, live demo when you got to do it on stage. It's real. >> Mic test. One, two, three. Are you there? >> Yeah. Thanks for the tip. Hello. Are you dead? Hello. >> So, right now I'm at the AI engineer conference conference >> and I'm like winging this live. I have really It's actually not the most set up.
15:10 >> Yeah, cuz like actually right now the speakers my mic. How was that? >> All right. So, one thing is that context is not just about best practices or external information. There's a lot of context in our heads. Whenever you talk to Codex, there's a lot of a lot of things that Codex can learn about you. And that's where memory comes in. It's really useful. It's I think about it as bringing context over time. So it is off by default but if you go to Codex app you can turn it on and let's say like for example I use UV and then Codex would learn that over time. I prefer using UV for Python package management.
15:59 But memories and context isn't just textural. It's also about what we're doing what's on screen and so forth. And we have a research preview called Chronicle for pro subscribers. Also worth checking out. And with that you could literally for example be on the GitLab GitHub looking at your pipeline it failed and you just ask Codex why is this failing and it sees you were looking at your GitLab your GitHub pipelines and then it'll sort of figure it out. So to sort of set that up it's quite straightforward you can just go to the Codex app you go into settings and then it's called personalization and you can enable it. So you can enable memories and also sort of given the as a safety feature you can also sort of skip generating memories in chats that involve tools because sometimes the tools might return information that you might not want persisted over time. So that's also worth considering. Right now I'm using a business plan so that isn't chronicle but if you're on a pro plan you would see chronicle as option to enable here as well. And then there's this thing called custom instructions.
16:56 So here you can sort of en it's like setting the agents.mmockdown in the home directory. It applies to all conversations. So for example, I said here always reply like a pirate. Then in the subsequent commentary, Codex would always reply like a pirate. A trivial example. So that's context. And to summarize, I'll let Charlie sort of summarize our first two dishes of the day. So to recap, part one, right, we want to give codecs context. you need to know which ingredients you're actually using for your recipe. that includes bringing in external content via plugins, right? Me personally, I am a pretty heavy user of Slack, Gmail, linear. and then it makes it really easy to just say like, hey, go look at my, you know, unread Slack messages for today and create a daily agenda for me to to think about when I get into work.
17:47 you want to be able to like keep the flow when you're adding your own individual context for the tasks that you're working on. I love dictation. I think just being able to, you know, as we've been demoing up here on stage, say, "Here's the context. go do it." I think dictation to me is a really underutilized feature because we speak so much faster than we can type, right? And these days, you know, we're we're beyond the the era where you have to make sure every single word is correct when talking to a large language model.
18:14 you can just give it reams and reams of your own rambling in transcription form and it'll generally figure out the right thing to do, right? And you're able to actually talk through the idea and and get to a better space by the time you send it to the model. likewise, appshots I think are are incredible and also very heavily underused. the ability to just like from any app I think I I don't copy and paste, you know, anything from Slack anymore. I just appshot my my Slack window or I just appshot my browser and send it directly into codeex and tell it go figure it out, right? And you know 90 plus% of the time it can figure out the right thing to do. Last but not least, we've got personalization. I think you know agents MD which hopefully folks know and love. and then memories as well in the Codex app via Chronicle. worth noting, I think memories, and Chronicle will use up a little bit more of your token budget, but, it can be a really magical experience to have the app just kind of understand what it should be doing based on conversations you've had it in the past. It really levels it up and makes it starts to feel like a true collaborator. there's a ton more stuff, I think, both in like context and personalization in the Codex app.
19:20 Strongly recommend you go check it out. pets is like a universally beloved personalization feature which we're not really going to showcase here today, but is would would yeah, definitely go make your pet. >> Cool. And once you have all the ingredients, it's you you're here to build and to really help you build and you want to distribute it. Sites is a great thing. At OpenAI, we've been using sites extensively. I for example have been using sites to summarize like an event like an event after like like after an event like today's I would ask hey Codex can you summarize what happened today create a site then we can share it with the team the rest of the organization I create a site to sort of we we were planning offsite and then Cory Cory on the team created a site to help us plan choose different locations understand the trade-offs so you can create artifacts to help you make decisions specific for code you can even you can also it helps with code review so let's just give a quick example with a live translate what we had just built.
20:17 You just do slashcode review and then you can review against uncommitted changes and then voila, you're firing off a code review task. It's a bit like having that you know sue chef like double check everything. You know, you've seen the movies right before the chef like sends out a dish, they always do a quick taste test. I think that's a bit like that. Image gen where you can create images. And let's see what the example Charlie kicked off created.
20:41 Oh, okay. It used SVG but not too bad. Not too bad. So then what you could also do is use the image gen skill and ask codeex to sort of cook it. So we'll come back to that in a bit but let's try that. Let's use this and spreadsheets document and slides. We won't be talking much about it today but you can sort of build these artifacts as well. And in the subsequent talks in the rest of the days and the AI engineers Jason for example will be sharing more about how he's sort of building with codecs for sort of these deliverables.
21:12 So those are the first two steps, context and building. But it doesn't just stop there. So if you're sort of new to Codex, it's your first time still using coding agents. I think that's really helpful. But I'm I'm pretty sure this AI engineers welfare. We're in a room full of power users. So for the next four steps, we hope we're going to share something that's practically useful for you. And the first one is computer use. So Codex can sort of take actions like as you see right now in the browser, Codex is debugging something. I won't tell you what's happening. I will we'll reveal the surprise at the end but sort of working in the background thinking about something. So let's give an example. So I I used to work in a large organization before open AAI and not everything had an API. You had a lot of enterprise software that had like they were legacy. There were different dashboards. You wanted to download data.
21:57 So let's say for example we have this dashboard that somehow pulls data from the OpenAI Python GitHub repo. There's this nice date range picker. But imagine I had to get data for every month from the start of the year the year to today I'll press Monday to Sunday, Monday to Sunday, Monday to Sunday, Monday to Sunday. It's a lot of buttons you have to press. Then there are different data sets you could choose. So what if Codex could help me with that? And the answer is yes. So again, I'm going to take a shot my two thumbs and then just take it. Okay, let's just make sure we're in a new conversation.
22:32 I'll just ask, >> hey, can you help me download the, number of commits for the last seven days as a CSV and the number of pull requests, the pull request data as a JSON from the start of the month to today. yeah, use like computer use and go for it. So again, it's like two different dates. I'm asking it for the like Monday to Sunday or the last seven dates days but also from like the 1st of June to today which is the 29th of June and you would see okay Codex has correctly invoked the computer use skill and very soon it's going to like start looking at the application you see here looked at application and very soon in the chain of thought summaries you would see more of it and very very soon again my hands are not on the keyboard and you can see this is my cursor I'm just going to like park it over here right next to enterprise is you would see another cursor pop up.
23:26 >> There it goes. >> And yeah, there it is. It's moving. It's downloading it again. My hands are here. Charlie's hands are here. My phone is like unlock. It's locked. And Codex is downloading all this data by itself. It's a bit like a mag. It's like pretty magical the first time you see it. And that's we're going to let Codex work in the background. And that's the beauty of computer use of Codeex. You don't have to wrestle with the agent for control of your screen. You can let it work in the background.
23:48 >> I I think it's also worth pointing out that on Mac OS, you know, it doesn't steal your cursor, right? Like Gabe still opening the app. He can still use a different tab. >> I can go scroll on Twitter and something. Yeah. >> Yeah. And so you can still drive your computer while things are just happening, you know, in the apps that Codex needs access to. >> So we're going to let that work. But sometimes not everything is like a native application. And there's some sometimes you're like maybe using Google Drive, you're creating a feedback form like for today's workshop, you know, I mean feedback forms are I personally find it very cumbersome to create a feedback form. you need to like create a question, you need to like fill in the different options and I always forget to mark it as mandatory even though I want it to be mandatory and it's you press so many buttons with creating these feedback forms. So I already have the questions drafted here and let's see if Codex can help me with that. So I just create a new thread.
24:38 So again, appshot, it's quite addictive, I must say, the user interface. And we just say, hey, can you create these this feedback form for me using the questions in the Google docs? remember to make mark each question as mandatory. And just to give codeex a bit of a I mean just to make things for exposition purposes, what we're using here is the Chrome extension, which is slightly similar to computer use. So that's this is computer use, but here it's directly controlling a specific application which is Chrome. So we're going to fire that off and we're going to see very soon another cursor come up. So I'm going to put this two side by side over here and you would see codeex thinking about the task. It realizes, hey, I have the question form questions already. Then it still needs to drive the live tab to actually create it. And again, like what we saw earlier, it's going to sort of start pressing buttons. So let's just wait for a bit to see the first button get pressed and then we'll swap back to the slides in a bit. So there it is.
25:40 You see it again. My cursor is here and then you see the blue ones here. So you can continue to go do different tasks. Let's say like maybe I'm interested to learn something about the codeex repo. We were talking about memories. So I can just like ask hey can I configure the model that's used for memories in codecs and what's the config term I need to change? So you see while you're continuing to do other tasks, codeex in the background is sort of doing the button pressing for you and we'll fire that off. So that's computer use and there are like different flavors of it like you know how maybe you know your favorite dishes come in different flavors and I'll let Charlie just really give a nice summary of what where the different flavors are.
26:24 >> Yeah, at this point there's there's like mentioned a few different ways that Codex can start controlling the applications that you're using. and it's kind of important to understand, you know, when you should reach for each one. The three that we currently have today are computer use, which you saw, the Chrome extension, which you also saw, and the inapp browser, which I think we did not demo. the inapp browser, why don't I start there, is a feature where in the Codex app, you know, it has it the ability to open, a browser tab on the side just like you would kind of a normal you know, desktop coding editor these days. but in addition to just opening and rendering web pages, Codex has really deep hooks into that browser engine deeper than it has at the Chrome extension. and so as a result, I think the inapp browser is really great for when you're building local web applications and you want to do a lot of deep testing on the interaction design on how it's rendering because Codex can see much further into the into the the layout at that point.
27:18 But going back to the two that we talked about today, we have computer use, right? Which is great for I mean anything on your on your laptop or your computer. but it really shines when you have applications that otherwise don't have a programmatic interface to them, right? it's great if it's a web app and you can reverse engineer the API calls, but you know there's a lot of desktop software out there. There's a lot of legacy software out there which doesn't easily expose that kind of interface to an agent. And so if you just need C codeex to drive the app in the same way that you a human being would computer use is really incredible and kind of unlike anything else that that I've tried it it to this day it's still just so magical to watch those little cursors fly around. then there's the Chrome extension which is your sign-in browser and I think that is the most useful for when you need to hand off state or authorization to the agent. So, I think we've all, you know, probably run into the issue where we want to build an agent to go and, you know, use the worldwide web the exact same way that we do, and the first thing that it immediately runs into is like, "Oops, I can't log into your Gmail account." and worse than that, like Google has detected I'm on a headless Chromium, and it's just sending me this infinite capture loop, right? and so the Chrome extension allows codeex with your permission to go and read from certain websites as you and borrow your credentials so that you don't have to worry about the authentication step as you're building.
28:38 Cool. so that's Codex take action and then next we're going to move on to collaborating on complex tasks. >> So this one's a really fun one. So I think people have really talked about how Codex is really good at like complex refactors, longunning tasks, and in the sort of next 15 to 20 minutes, I want to share some tips on how you can sort of work best with codecs on these complex tasks. So, it's going to be a bit hard to like kick off a 20our longrunning task and we're all just going to sit here, enjoy lunch, dinner, and breakfast. So, I started off some tasks the night before, and they're like four different tasks, each covering a slightly different flavor. So, this one, it's still running for like 15 hours.
29:20 It's running on my remote instance. >> You can see at the bottom here, I'm running a goal. It's been running for 15 hours, 30 minutes, and 55 seconds. So, I got to scroll, scroll, scroll, and here I am right at the top. So, it's a bit of like a green field project. I want to build a live Q&A site. it's going to be in a room for like 400 people. There's like a admin view. There's a participant view. There is a onstage view. And then like I describe, oh, maybe the admin can do like AI generated tags. There's a moderation features on the stage view. There's a 16x9 run every application through our moderation API. make the experience feel like something OpenAI would launch so on and so forth. here I'm being particularly prosaic in the prompt for exposition purposes. You could just ramble to codeex as well and it'll get it'll get the job done. But here just like exposition purposes, I know people are excited to take photos. This is like an example of how you're being clear. I think one misconception these days is like people always ask does prompt engineering matter? In so far as the instructions are clear, it still matters. Like I can't read your mind.
30:27 Codex can't read your mind. So you just need to tell Codex very clearly what you want. But it's prompt engineering where you see as oh I got to use this magic phrase. I don't think we're in that state anymore. Reasoning models are really intelligent and they sort of the real thing is about context. So actually goes back to our first step giving context clearly. So how I wrote this prompt was I was just talking to Codex.
30:45 I was I was yapping it on yapping to it on the phone. I just hey Codex can like polish this prompt so I can show it on stage. So yeah it's building this thing and I'm using like a few different things like convex. There's a a plugin for convex that helps to do real-time interactions going to deploy in a cell and of course the open AAI plugin and all right the feedback form is done. So that very nicely came up and yeah that's one example so it's like a green field project but you know not everything is like you're going to start from scratch.
31:13 So here's another example that ran for 5 hours and 57 seconds. So here's an open source repo for a form builder and I wondered, hey, could I like repurpose this and build it internally in sites? So maybe you don't want to buy SAS. It's a small team of like seven of you. You want to just have a internal form builder. You want to use the chat GPT authentication that comes with these sites to sort of ensure that only people within your team use it and then you got to let Codex sort of build it. And I I will show you an example what that looks like later. Then the last example of a long-running task that was running. This one ran for let's see I think 7 hours 1 minute and 40 seconds. So it's like some gradu research papers from my graduate school days. it was like written in like Julia then I think like an R and I was like hey can I like write a Python package with a Russ backend and then I told Codex again I gave some constraints I don't want to use dynamic programming you got to use like some heruristics and then codex read the papers and did it. So those are the three longrunning tasks and I want to like give some examples and some of the best practices that come with it. So in one of these for example what I told Codex was this over here you can define a goal. So there is actually a skill in the openi skills repo called define goal and I told codex hey don't create one goal create a series of goals and put it in a goals.mmuckdown file. Again, it's just very arbitrary, but it's just a collection of files that you sort of see and you can can inspect it and then create a dashboard, progress dashboard.html, so you can go back to it and see what's going on. Because in this seven hours, you don't really know what's going on. You want to have some some way to like see what's happening.
32:51 And rather than having this one thread do all the work, it's actually spawning new threads. So, if we go to the see, I've pinned the project, but if we go to the the thread, but if we go to the project, you'll see it's it's spun many other threads. So for example at different milestones it ran code review and then it also ran a goal audit. So it's seeing like hey are you on track there's this initial thing you talked about like goals mount down are you deviating from it? If it is then it will nudge back the main thread. Hey remember this or do we need to rethink the plan and so forth.
33:23 So it ran at different times. That's why there's like different milestones M2 M3A and so forth. there's another example here with like the live Q&A site. this is still running. So this is running on a remote session. So it's like my gaps remote and I want to like do testing. So what I told Codex is that hey you can spawn a local thread on my machine and use the Chrome extension to then inspect it. So you would see for example this is a thread that came and you would see it says here send by codeex from another thread. So the remote thread can communicate with the local thread and says hey you got to do click ops you got to do browser you got to do QA testing. Then the local thread did all the QA testing and it send the results back to the remote thread. So that's another thing that's worth thinking about for these longunning tasks. And let's see what else we can call out here in the long running tasks.
34:12 communication. So one other way of thinking about this is that maybe you could just ask codeex to send you Slack messages. So here I have a few demos. So, I was building a live Q&A site and without having to set up additional infrastructure, you can just use the Slack plugin and just say, "Hey, reply to this thread and give me like a summary of what's done, what's not yet done, and are there any blockers?" Because sometimes there are blockers.
34:36 Maybe there's like a authentication issue, there's an API key that you need to give to Codeex and so forth. So, I asked Codex to speak to me. And over the weekend, like on Sunday while I was out with my friends, I was just regularly checking, is there anything I can help unblock Codex for these different tasks? let's just give some screenshots of what that process looked like. So just take a look at it. So another nice feature with codeex is that while you have the main thread. So right here in you see in the screenshot on the right left is the main thread on the right you can start a site thread.
35:09 So you literally see here and say let's focus on signing in with chat GPD first. I want to sleep soon. I assume you need me to do that because I know like okay I think Codex needs probably some help with that because I'm running a building a new instance on the app server a new user interface so it needs to get some credentials so it needs to know where my credentials are and it says okay actually the credentials are already here so okay update the main thread so you can use the site thread to update the main thread so let's do an example of that live as well so here this one is still running you could just do slash site and it creates a site thread and you can just ask codeex like what's going on. So maybe we just ask it takes a while this I think due to the network but let's say hey what's going on can you list like what you've done so far maybe in the last hour what do you plan to do next is there anything I can do to unblock you and then from this when you get the agent tells what's going to do next you can also get a sense of like what maybe it's like on the wrong wrong track then you can sort of tell codex hey maybe you don't want to do that we should do this instead then the site thread can communicate with the main thread so that's a pretty nice feature >> it's worth noting with side threads.
36:13 Those are ephemeral. So if there is any context that you want to keep longterm in your project, make sure you're communicating that back to the main thread because they will get stale after a while. >> here's another example where what we were doing earlier was a site thread talking to the main thread. This example is where two different main threads are talking to each other. Specifically on the right the left side we have codecs saying test spawning a local thread in the local form builder project. If it succeeds make that the default location for all local future local threats especially for Chrome or end to end testing. So again on the what's happening here is that it's running on the remote service but so it doesn't have access to my chrome chrome browser. Then it send a message to the local host and then it can then from there see chrome appears to be installed and then from there it can do the testing with chrome. So there you can get the remote session to work with the local session and this was the test that it did. So while it's working in the cloud locally on my laptop like me at 3:00 a.m. while sleeping Codex opened Chrome it went through the form building flow and it created a form end to end.
37:23 Sub aents are also a particularly helpful thing in these longunning tasks because you want to make sure context is well separated. agents sort of focus on different tasks and here's one nice example. So it was building a application and I just like took a peek at it and I just not necessarily best practice. I just sent it to the main thread saying not very good. Yeah. And then Codex said yeah agreed that's technically cleaner but it's still a bad chat interface. And you notice here what it did was that it messaged the sub agent. So the sub agent was already previously spawned but then the main agent messaged the sub aent and you would see this the message over here. So for example it says this. So it sent it and says update your verification task with this new user screenshot and directive. So there you have the sort of inter agent communication.
38:11 so this was that what it did thereafter. So it did its own tests because I sent a screenshot with the messages flushed to the bottom. I didn't want that. So then Codex did its own test and it says send retroex visible okay in one short sentence and it sent it. So CEX itself did its own QA testing to sort of verify my feedback. and this is an example for the live QA where the main thread was spawning different other site threads for example the audit of the goal the code review but also doing that communication back to Slack. So this thread over here where it says provide slack thread URL that's over here where I was providing all these updates. So a separate sort of session was managing that.
38:52 We earlier talked about goal and here is a dashboard that the codex created for the goal. So you would see it has all these different milestones, what's been completed, what's active, what's yet to be started. And then the dashboard also gave me a sense of like what the different threads are doing like oh this thread is the main orchestrator. There's a communication thread. There's one thread doing like visual UX. There's one doing like the managing the app server protocol and so forth. And even like creates it into like more conceptual work streams like one workstream on off one workstream on testing one workstream on UI polish.
39:28 But throughout the night sometimes codeex does get blocked. So then with the example of convex in versel, it said oh I I marked it as blocked because the critical path it didn't have the authentication didn't have like the identity to access these external services but still I did everything I could up till that point. So it said it was blocked so it will come up in the UI and then I just selected it and I say add to site chat and it says can you please explain what's going on because I just woke up like I I didn't have context of what's going on but Codex is telling me he was blocked. So I'm trying to get Codex to to catch me up to speed.
40:01 So here it's like saying explain step by step what I can do to unblock you here. So it gave me different options how I can give the code VLE details the convex details and then I sort of did it. So here it I sort of I selected it and say do that. So then the site thread told the main thread to sort of create the CLI login flow and I'll just click it over here and I'll be able to then authenticate and give vers codeex which again is running on a remote instance access and the right permissions to manage my versel project.
40:31 Then in the case of convex there's a deployment key. So I was like hey I'm not going to paste that into the chat. So I asked Codex give me a secure way to do it. Then it gave me this like bash code that I could just paste it in and I could then set the key into the remote environment. So that was like a way where you could collaborate with Codex in a remote instance, do different testing, managing permissions and so forth. so let's see what they actually built from this entire process.
40:56 So let's look at a site builder site one. So this is it, the internal form builder. So to recap again, there's an existing repo where of a form builder, but I want to like customize it for like maybe my internal team. We're like maybe five of us. I don't want to spend money on SAS. So we can build our own form builder. So it creates a new draft. You could call it like AI engineers workshop. Hello. And then just say like yeah I should I should ask the Chrome extension to do this but you know for exposition purposes.
41:27 Okay. And you can like preview the form and then you can sort of share it thereafter. You can like copy publish the link. Then there'll be a sharable link. You open the link and then someone else in your team can fill it in. And again because this is hosted on sites only people that have access to your chat GPT workspace can then access this form. So that's really nice. Yeah. Then you can see all the responses and so forth. So this was again what we did. So to summarize on the complex task, we had three different projects. One was a machine learning one where I wanted to reimplement a package with a Rust back end. one was to take a existing to build from scratch like a green field project a live Q&A site and the third was to take an existing open source project and customize it for my sort of needs and codex was able to sort of get it done in those cases. So underlying this are a few sort of concepts and primitives here. Let's go back to the slides.
42:28 So the first thing is compaction. You saw how codex worked for seven hours. It worked for 5 hours. It worked for like right now I think the other one is like 12 hours. And many times it's compacting its context window. This is not a new technology. In November 19 when we released GPT 5.1 Codex Max, yes the name was GPT 5.1 Codex Max, we natively train the model on multiple context windows through a process called compaction. And that's what lets Codex be really good at it. some examples of what people have said about compaction to quote did open AI basically solve compaction I pretty much never had issue with 5.5 and codecs across ultra long threads spanning many compactions. this individual from Japan said like in fact codeex has strong compaction. So even if you make a lot a lot in the same thread it's less likely to forget. And in that sense because it only remembers the good stuff. There's some token efficiency to that I suppose. And as always, people think it's pretty quoted.
43:29 I'll let Charlie share a bit more about like specific primitives and concepts that enable this such as goal. >> Yeah, thanks Gabe. So, as you saw, we we have a goal that's already been running for 15 hours. And I think you know goal is is one of these primitives that is incredibly powerful if you can use it in the right way. one of the biggest things that I tell people who are just trying it out for the first time is be really specific with the criteria that you're giving the model to to know when it's done. and specifically, you want that criteria to be as verifiable as possible.
44:02 sometimes I do feel a little bit like goal is is kind of like a genie in the sense that when it works well, it is insanely magical. and when it doesn't sometimes I feel like I've entered into a monkeykey's paw type situation where you know it has done the thing technically that I asked it to do but like very much not in the way that I was expecting. so it when you give it a task it'll run off it'll keep evaluating itself against the success criteria against the verifiable criteria that you've given it. and at each turn it'll check hey did this actually complete and if not should I keep going right there are quite a lot of use cases for very large projects in this way. you know, Tama, as you can see, used a SLGO that was running for 40 hours, that did a reimplementation of Doom in native Swift code. if that's the type of thing that that you want to give a shot. but also things like code migrations, large refactors, retrying loops, experiments, and, you know, like games or or full oneshotting apps with fairly detailed specs. I think goal is is pretty incredible.
45:04 we've also got sub aents which we saw a little bit of a preview of here right sub aents in codeex you know by default will get used out of the box. You can tell tell the model hey you know think about delegating to sub agents here but you can also customize them. So if you want if you know you have a sub agent that should be your you know code reviewer right you can have a specific prompt set of developer instructions here. if you know you want that code reviewer to just be focused on like just be a lightweight model so that it runs quickly. Or if you want it to be a much more intelligent model so that it's very thorough, you can customize that in its own sub aents.totml file. and you can give these names. You can say, oh, go, you know, check with the code reviewer, go check with the docs researcher, in order to to implement, right? Sub aents can work in tandem with thread-to-thread handoff in codec. So, thread to thread is is quite new. I think we just released it in the last week. Yeah.
46:02 and so that's just an even higher level of abstraction. and the way that I think about balancing them is, when the the delegation of work needs to be visible to you, the user versus when it just needs to be visible to the model. So sub aents we started with because I think that is when the work, you know, primarily needs to be visible to the model. if you want to delegate, you know, some like a subcontractor, you know, I don't really care what they're talking about. I just want to make sure that the work comes back good. Sub aents are a great fit and the model is great at figuring out when it should be handing off into sub aents. it also keeps the the context separate. I think unlike you know a lot of other tools that we've we've brought to codeex to help keep context windows efficient. sub agents are one where you know, you're keeping the context separate and the mo the the models are communicating with each other. Thread to thread handoff is useful for when you still want to be in the loop and you still want to be able to like look at everything that's happening. and perhaps even larger than that, you are mentally thinking about separating these two these two threads, right? So, one example is over the weekend I was working on a game and I had a creative director thread, right? I could have used a sub agent and that would have worked fine, but I wanted to see all of the different things that were happening and I I didn't want it to just be delegated into the ether via sub agents. So, with the creative director thread, I opened up different threads, one to work on the art direction, one to work on the music, one to polish the animations, one to figure out the game mechanics, and I was using those and I told them in the agents MD file, when you're done, go and check with the creative director thread, which I'd given a bunch of extra context on, you know, the look and feel I wanted for the game. Go check with the creative director and get sign off before you consider this task done, right? and they all did so. that, you know, it was actually really interesting to watch the threads just constantly sending messages to each other back and forth as they worked on the game.
47:56 we didn't quite show it here, but we also have hooks in the Codex app. if you want to set up your own hooks, you can do so in your config.toml file in your codeex home directory. these are really great for when you want to start introducing deterministic behavior at specific checkpoints in your software development life cycle or in your just you know general development life cycle. So I think in these cases you might use hooks to do things like you know send your conversation to a logging engine right you might want to scan inputs or outputs. I think there's quite a lot of security use cases for hooks, right?
48:31 Before Codeex takes a specific action, maybe you want to look at the prompts and figure out, hey, have we accidentally pasted an API key that's that's going to go somewhere or before it takes an action to call a tool? Hey, is this tool, you know, on like not on our approved white list or something that we need to be extra sensitive about when calling the tool to evaluate if it's if it's safe to use. and we can also do custom validation checks as well when the turn stops.
48:55 these days, you know, you can tell the agency MD, hey, go and use the the llinter, right, before you consider the work done to check your work. and it's really good about doing that, but if there are things that are a little bit more sophisticated or a little bit more complex that you want to do to validate the output, you can use hooks to to take care of that as well. you know, I find that when building more and more complex projects, some of the most important things are figuring out, like I've said, the verification criteria or the success criteria. You want to give the model as many boundaries as it can as you can and let it fill in the lines in accordance with like the boundaries that you've given it.
49:31 >> Cool. and that brings us to working over time. >> Thanks Charlie. So just to quickly share on hooks slipped to my mind as well. So we have the OpenAI agents SDK and there's actually a quick example of hooks right over there. So if you go to settings, all the hooks are sort of there. And you would see the hook for the agents SDK repo. And here it's sort of doing at the every end of its turn a Python script to just tidy the repo. So again, there like some of these housekeeping that maybe you want the language model to do. You can run it deterministically with a hook. And hooks become particularly useful with longunning tasks because there is that trade-off. You're giving the agent more autonomy. It's going to work off work longer and longer without your supervision or abstracted supervision. And to that to that end hooks become useful in getting the guardrail.
50:22 So on working over time what we're trying to show here is that now we sort of collaborated codecs we have some of the power tool power user tools like computer use. How can we like sitch this all together and let you know codeex work on its own? I mean for exposition purpose here we are like prompting and like being more methodical but in reality you can just like let codeex work on it like with thread handoffs. In practice, I could have all my demos in one thread and just ask that thread to segregate to separate it and send it and delegate it to the various threads. So here's one quick example of a automation. So right now you notice like on my screen over here that these like blue orbs. So this is just showing that it's a remote host and maybe you want to create one. So we have a nice plugin here with the digital ocean plugin which we just released last week and then you saw do it. You can just try it in chat. I'm just going to like fire it off as well. And what it's going to do is going to like provision this infrastructure. It's going to take some time. And I'm not going to like stand here and look for for it to provision this infrastructure. and Codex is not going to like just keep like waiting. What it's going to do is create a heartbeat automation. So it's going to like monitor and so forth. So it says yes. And let's just show in the spirit of this being a cooking show an example that was already done. So here provision a digital ocean droplet. Okay.
51:43 Yes. And then it created the automation here. I set a codeex heartbeat to resume this track in about 5 minutes. So it's going to keep tracking every 5 minutes. Is it ready? Is it ready? Is it ready? Is it ready? And once it's ready, it will take the next step. So in this case, it did it. So okay, the droplet's ready. And it even gave a nice link. So if you click this link, it's a deep link. It's a link like a deep link to the settings page within the Codex app which automatically sets up the SSH instance for you. So you just need to press it and then it'll preconfigure everything. It knows where it will create the SSH key for you if necessary and so forth. So that's nice and you can extend this to other parts of like say the software development life cycle.
52:24 maybe you're building something you need to deploy. You have a pipeline that triggers event based and it takes time. So you can get codeex to like run every 5 minutes, every 10 minutes check is it ready, is it ready, is it ready and maybe depending on the event then it sort of spawns different tasks and so forth. So that's a heartbeat automation. It runs within the thread. But we can also do automations that sort of spawn new threads. so one example of this so right now I have actually the Slack channel. So see if it's up. So remember the Mac OS translation app. So in the last say 45 minutes we had a bunch of users all called Gabriela and they've been using it. They've been giving a lot of feedback on the translation app. for example, oh the screen feels focused. I can tell where to start, so it's positive. Is there a way to do a quick sound check before the audience sees caption? So, so there's a lot of feedback. People are giving feedback.
53:20 And I want to like, you know, give this feedback the attention it deserves. So, I want to create an automation to do that. So, just make sure we are in I'm just going to close this since we don't need it already. Over here, we create a new thread. And then I'm just going to like do a appshot. every half an hour, can you just like look at the comments and suggestions in this channel and reply to them? I think there'll be like a few categories. There's one category where it's like compliments. There's one category where it's a bug report.
53:54 There's one category where it's a feature request. so forth. Let me just like fire that off. So, it's going to create an automation. And I I just sent it, but I realized, oh, wait. if it's a feature request or a bug fix, I want to do something different. So, I'm going to like steer this. Oh. if it's a bug fix or feature request, can you also then like spin a new thread in a work tree and actually address that and then open a PR and then like request Dom to review it and then so forth and ask Dom to like approve it within like 24 hours. If not, just keep reminding him. So yeah, it's it's it's a comical example, but here it's where I'm try I what I was first doing was steering. So I'm going to I don't have to wait for codeex to finish. I can sort of steer it midway. And what I'm trying to do is in addition to all this feedback and classifying it if it's a it's a bug report, if it's something that we can fix. Codex would then sort of spawn a thread. So we're putting all these primitives together, the automation, the plug-in which can can read slack. The fact that one codex thread can communicate with other codex threads and even create work tree threads and so forth. So imagine then you wake up and this is a bit like that software factory where all the PRs are ready then you can even add additional comp layers to it where it's all this review you build a test version you send it to someone they test it and then they give the thumbs up then it's good to go that feedback gets fed back to the PR then the PR reviewer sees all this holistically. So yeah, there's a lot of it you can chain together to build that automation. So then every half an hour it would do this and it will find DOM as well in the channel. So you we were talking earlier about how we've actually been running a automation in the background and let's find that. So let's see here it is. So earlier while everyone was coming into the room I was asking Codex can you run an automation every five minutes taking a screenshot of what's happening on screen the Google Chrome Google drive presentation we only have 16 minutes and I'm asking Codex to estimate to what extent are we going to overrun so right now it's at checkpoint 13 it's seeing if the demo is there it's flashing 12 10 12 and so forth so Codex is doing its own assessment of like to what extent are we going to finish on time so it's a bit of a meta example of how you can use automations and chain together other tools like in this case the Chrome extension.
56:12 >> It looks like it's telling us we might be a little bit behind the ball here. >> Yeah, it's it's a you know it is what it is. So automations have been really useful in my personal life to extent really like getting a lot of things done. some automations that may be useful too. one automation you can run maybe every Friday you ask Codex to look at all your conversation threads and see are there skills that you could improve? Are there like things you could add to the agents markdown? Or are there skills you could just remove? maybe skills that you haven't been using. A second automation or approach that may be useful. Maybe you're using automations to draft email replies. You can have one automation that drafts the email reply, but you can run a second automation that sort of cross validates against the eventual reply you sent because that is the ground truth. Then it updates like a markdown file with the best practices, tips and so forth. And in that regard, the email drafting automation can get better over time. So, I'll let Charlie just recap what automations are and share a bit more about the app server as well. Our final step.
57:14 Cool. so I think yeah, automations there's two types two types of automations, right? I think the more powerful one these days tends to be the inapp or inthread heartbeat automations. And so those just live in a single thread. They run on a timer. they you know keep going over and over until they're they're done. these I find like have just almost unlimited use cases. me person, Gabe shared some of his. I think me personally, I feel like I'm drowning in Slack and email and linear notifications every day. so I have an automation that just sort of, goes through and checks the latest updates from, you know, all the sources that might need my attention. I have a local Obsidian vault and so it's just constantly updating my daily notes. It's constantly updating my notes on different work streams and it's keeping a lot of context there. And that way I can just pull from that later if I say, "Hey, you know, what's going on with the latest AI engineer world's fair talk?" Like, you know, how far are we on the demos, how far are we on the notes? And then I can just Codex can just quickly tell me, here's what's happened since the last time you checked in. but automations can also live in a new thread. If you set them up in the Codex app, under scheduled, the scheduled page, you can have them spin out in their own thread as well. and then that's useful if you just want to say, "Look, every week I want you to do a recap of what all of my direct reports have been doing, or I want you to look at all of the analytics from this system, and then we're just going to do a one-time look together, and then I'm going to archive the thread and not worry about it again."
58:37 I think last but not least, we've got embedding codecs anywhere. and so I think we're for this one, I want to talk a little bit about the app server protocol, right? Which is, the core protocol that underpins the Codex app, the Codex CLI, the Codex VS code extension. the app server is basically, what allows all of these services to talk with the main codeex harness, right? and it's integrating all of the tool calls and the plugins and the context compaction that you know and love. we have you know built a lot of the components of the Codex ecosystem to be open source because we really want to encourage developers to build wherever you are and for whatever use case that that you have.
59:19 And so to that extent you know if you're not aware you can embed codecs into your product using app server. I think this brings a lot of the power of codecs directly to you. and it brings the ability for your users to use their chat GBT subscription and perhaps more importantly use their chat GBT token budget to power the products that you're building. and like I said, it has all the same underlying features that we've been demoing. You know, compaction, steering, the ones that need to live on your computer like computer use are are, you know, not going to be there. But for everything that exists in the core, you know, agent harness, we've got that in the app server.
59:58 So as a quick demo of that, so there was a fourth longrunning task that I kicked off last night. It was actually to build my own user interface for codeex with the app server called Retroex. So it's been working for a while. I think there are different tasks I worked with it. And this is sort of the work in progress retroex. It's using the app server. let's just okay zoom in then optimize it for the stage. But yeah, you can like choose the model 5.5.4 mini the reasoning effort. So I told like oh make it a bit funky when you change the reasoning effort. So at low there's like a small glitter at medium the color changes you think high and extra high it's like woo a bit more. So you could just like simply say like hello and let's see this is in retroex going to send hello and then allow access. It's going to run and it's going to reply back hello. And what's nice and to just prove that it's all running on the same app server is that if we go to the Codex app in the same project retroex you should see that over here let's see yeah it should we can change the threads I believe but yeah that's an example of how you can build it's it's a trivial example of building a funky user interface but let's say you want to build agents on the cloud you want to manage different agent providers you're considering in a large organization your different sort of providers you're using and you want to build a unified layer the app server is a useful way for you to integrate it into your own platforms and because it's open source if you go to you get clone openai/codex you can just ask codeex about codeex you could just say like hey can you tell me more about the codeex app server and then you can fire it off and you can start building your own applications or integrations from there immediately so that's the app server And we have a nice blog post talking about it and there have been more engineering blog posts about like our Windows sandbox first of its kind how we built the agent loop amazing stuff and it's a good reference point also like if you're AI engineer seeing the codeex repo it's a great source and we'll have further talks in the subsequent days about this. So I'll let Charlie share some closing thoughts after seeing this entire cooking show.
62:13 thanks Gabe and thanks everybody again for being here. I think hopefully you've learned a bit today about how you can use Codex to kind of level up your workflows. Some of the co closing thoughts that I want to end on here are, you know, questions that I repeatedly ask myself as I'm trying to get to the frontier of what Codex can do, right? they start pretty simple just by asking, have I tried asking Codex to do it? Right? I think like one of the things around the office that we say more frequently than ever now is have you asked Codex? but I think as you evolve that workflow, right, you want to go from just asking in individual turns to building loops, right? We've all heard about loop maxing. That is the new thing. and I think we want to think about how does that how does that work, right? Like what is it that the loop needs to be checking at each turn, right? as you go down this path, there's going to be additional questions like, okay, you know, as it's progressing, how how like what blocked Codeex from actually getting to the answer, right? was there context that it needed? Were there permissions that it needed? Are there ways to build that context into repeatable tools like skills or plugins that Codex can use the next time it runs into this problem? how do I start scaling kind of from one agent to many?
63:21 Right? Like as you think about the work, is there ways that you could be prompting Codex to implicitly delegate the structure of the things that you are building? or to start orchestrating across multiple threads to have different Codex instances communicating with each other. And then ultimately you know if you take a step back and ask the question okay why does this process need a human at all right I think like as we've shown you know one thing that you might assume when building with codeex or when building with the open AI API is that a human needs to be involved to like go to the platform page and go get an API key and then paste it into the environment right I think like as we've shown with the developers plugin that is not the case anymore the plug-in itself can just go fetch the API key and set it up locally for you and ultimately you know we do still want humans in the loop like where they need to be but for everything that's not that like how do we figure out how to take the human out of the process to unblock ourselves and how do we figure out how to build the machine that builds the machine right I think like it is one thing to sit with a coding agent and build a single piece of software is another thing to start zooming out and say actually I want to build a software factory right I want to build a system that that then builds individual pieces of software on an assembly line as we and so a lot of these things are kind of how I start to get from, just an individual ask to something significantly bigger.
64:43 >> Cool. And I think with that, we are >> a recap of >> Yeah. Yeah. >> Good. So, thanks so much for joining us today on a 9:00 a.m. on a Monday. It's the I hope it's a great way to kick off your, AI engineer welfare. And that's not it. So, tomorrow we have our opening keynote. It's going to be fun. Do check it out. We also have Jason on the developer experience team giving a talk about how he gets the most out of Codeex. It's a nice foil to today's presentation about how he's using on a day-to-day basis. And then he has a workshop after that where he goes through step by step like setting up the plugins, doing app shots and so forth.
65:22 So today's session is more like a whirlwind tour. This workshop will go through the step by steps of how to get to set yourself up for success with codecs. We then have a session on from the harness about the harness and then sessions about voice agents. Charlie will be talking about that as well. Dom will be doing a session on building on the Codex harness and then another session explaining behind the Codex harness. So the first one is you build on top of it. Then the second one is explain the harness. Then lastly a session about LLM inference in production and so how we sort of do that at OpenAI.
66:01 So thank you for joining us and that's the end of the first segment of today's workshop where we sort of give that whirlwind tour of everything codex and hopefully you have something fun to cope with. So it's time to build and to help you with that for today's session. first things first, if you've not downloaded downloaded the Codex app, do give it a download. And for credits, so we are giving 100 USD in Codex credits and 100 USD in API credits. So we'll leave leave this QR code up there's their phones up and take a photo. Yeah, I could like appshot this and say, "All right, send this via email to everyone who attended, but we don't have your emails." Yeah.
66:38 >> Cool. And I would also mention please stick around to the end of the build session if you can. We brought swag for everybody. So at the end of the 2our mark on your way out you can pick up a sweatshirt I believe. >> We love to see what you're building as well. Yeah. >> Yeah. Well, we will be here. We will be here in the audience. We've got opening eye staff here as well to help unblock you or to answer any questions about using codecs.
67:03 >> Thank you.
Summary
- Codex can be used for a wide range of applications, from building plugins for music software to creating self-driving golf carts.
- The workshop introduced six key steps for effectively using Codex, including providing context and building applications.
- New features include mobile accessibility, record and replay capabilities, and enhanced plugins for seamless integration.
- Codex can perform complex tasks autonomously, utilizing features like computer use, which allows it to interact with applications on behalf of the user.
- Automations can be set up to run tasks periodically, enabling Codex to manage ongoing projects without constant supervision.
- The app server protocol allows developers to embed Codex into their own products, leveraging its capabilities in custom applications.
- Personalization features, including memory and custom instructions, enhance Codex's ability to adapt to individual user preferences and workflows.
- The session concluded with a call to action for attendees to explore Codex further and participate in a mini hackathon to apply what they learned.
Questions Answered
What is the Cooking with Codeex workshop about?
The workshop introduces attendees to the capabilities of Codeex, showcasing various projects and applications that participants have created using the tool.
How can Codeex assist in building applications?
The speaker shares their experience using Codeex to build a Mac OS app, emphasizing its ability to provide best practices and context, even for those unfamiliar with certain technologies.
What advantages does Codeex offer for web application development?
Codeex provides deep integration with browser engines, making it particularly effective for testing local web applications and interacting with software that lacks a programmatic interface.
How can Codeex facilitate collaboration in project development?
The speaker describes using Codeex to manage permissions and collaborate on projects, including creating a customized form builder for internal use.
How can Codeex streamline feedback and development workflows?
The speaker illustrates how Codeex can automate the handling of feedback from various sources, categorizing it and initiating appropriate actions based on the type of feedback received.