Please welcome to the stage member of technical staff at Anthropic, Puneet Shah. Welcome to the last session of Code with Claude here in London. Yeah, let's give it up for everyone who's come before. So I'm a product manager on our team, on the platform team at Anthropic, and I've shipped some of the features like our 1 million context window, fast mode, and a number of our improvements on prompt caching amongst others. And that's what makes me especially excited about this session. You guys have made a very good decision good decision to come here because this is about the cloud platform. And if you think about what we have, we have these great models, but there's this whole layer on top of it, the platform, that's not about just getting you the intelligence, but helping you build a real business, a real product that really delivers for your users on top of those models. And so in the spirit of this being the last session, I want to like, let's get a little movement. If you have built an agent of any sort, big, small, I don't care what, go ahead and stand up. Okay, good. Nice. A lot of agent builders here. This is nice. Now, if you have put one into production, go ahead and stay standing and everyone else, just take a seat. If you're really happy about the quality, about the cost, about the speed of those agents, stay standing and everyone else take a seat. Okay, nice. Nice. Okay, look around. Remember these folks, these are our true experts. Go find them at happy hour afterwards. You can all have a seat now. Thank you. Thank you for humoring me. What I want to do in this session is share some of what we've learned about what helps you get the most out of that cloud platform. And one of the first things that we have on the platform that I want to talk through is prompt caching. If you remember nothing else from this session, think about prompt caching. And what caching is, is it's a way that we, if you're not familiar, where we take your input tokens, we process them, and then we cache that before we generate the output tokens. And that cache, we then continue to reuse as the conversation moves on. And so when you have a new message in the conversation, we just process those additional new tokens, but the rest we just pull from the cache. And okay, why is this useful for you? Why should you care? Well, the first reason is you get a 90% discount because we're not reprocessing, then we pass those savings on to you. You get a 90% discount, so huge cost savings to actually build your agent. The second is that you also get a rate limit boost effectively. If our rate limits are not, they don't count your cashed tokens. And so if you have an 80% cash hit rate, meaning 80% of your tokens are cashed, then you effectively have a five times larger rate limit in practice. So that's great. And then the last benefit is latency. If you are starting to cash a lot and that conversation is getting a lot longer, because we're no longer processing all those tokens, what ends up happening is your time to first token goes down. And so these are great benefits. And if you're kind of looking for a target and you're building agent applications, something in the kind of 80% above range is a good place to try to target. But if you look at some of these customers here, we have Replet, Cursor, Perplexity, Cloud Code, they're hitting 90 plus percent. These people have really talked to all these customers. They put a ton of effort into making their prompt caching work because of all of those benefits I've described. And one of the things that they all start with is just understanding what is your prompt cache hit rate. What is the prompt cache hit rate you get? That's the first question you should understand. And thankfully, because we've learned how much work these people are putting in, we said, why don't we build it for you guys? And so now today, if you go to the console on the cloud platform, console.anthropic.com, you can actually see analytics right next to your cost and usage pages. You can actually see analytics on prompt caching. And even just yesterday, we continue to improve this. Yesterday, we just launched ways to actually figure out why has your cache broken. Turns out the ordering of how you make those prompts matters a lot. A common error I see is that people put a timestamp into the system prompt. What day is it? OK, that's useful. But then it breaks the system prompt because it changes as you go. And that then breaks your cache. The tokens need to be exactly the same. And so you can actually see how did it break. And you'll see that on the analytics page. And that's a great place to get started. If you're seeing it at 0%, it's okay, that's why you're here. You can start with a one-line code change with auto caching, implements kind of a basic prompt caching, or better yet, go to the Quad API skill. Go to Quad Code or many other coding agents, and we have a skill built-in, prepackaged, that lets you ask it to improve your cache hit rate, and it will help with how you manage and order that prompt to get the best performance. So prompt caching, super important. We're going to talk a little bit more about it, but I want to talk about something I've been working on. I've been really excited about it. You guys are here for my start-up pitch. Thank you so much. I'm also in my free time the CEO of HeroCorp. And I have my CTO actually here, Ben. Do you want to come out here? He runs our technology team. Let's give it up for Ben. And can we switch over to the demo laptop? So we're, wait. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. This is a good one. We are a superhero for higher company. We help fight crime in the city, keep the tube running on time, staff your three-year-old's birthday party and make sure there's sufficient number of balloons. We do it all. And some people have misconceptions, I think, about the superhero world. They think that we're like, fly by your seat of your pants. We're analytical. We use OKRs. We plan. And that's why we have a dashboard. And this dashboard is how we track our performance. You can see 90 day retention, 94%. We could do better. We got one flight risk there, probably because our compensation is a little bit low. Okay, yeah, I probably should be paying the superheroes more. You can see some of our different superheroes in the Anthropics cinematic universe. Our lawyers required us to use these names instead of others. Are listed here with their quotes about how things are going. So do you mind pulling open our, we built a little developer console. You'll hear me repeat this a number of times. You should always look at the transcript of your agents to really understand what's going on. And so right here you can see we're pulling in a bunch of data and if maybe you wanna pop open one of those, you'll see here there's a bunch of data we're pulling in from web, from Slack, from Gong, from Jira, all the different sources to aggregate it to take a big holistic view on what's going on at the company. And what's our prompt cache hit rate? You want to take a look? Zero, then. Come on, come on. OK, let's implement prompt caching. So that's the starting point. We know what our prompt cache hit rate. It's OK. It's at zero. We can do better. We're going to get there. And so right now, at about 31 pounds, you saw it before it flashed away. And we've just implemented prompt caching. And what's happening on the right is the exact same output. Remember, prompt caching has no impact on intelligence. It is the exact same thing, just pre-processed and saving you money. And now instead of that long block that he showed you earlier, we're just saving that in cache and then reusing that. And if you look there, if you kind of zoom in a bit, you'll see that it says cache write 172 tokens right up there. And then cache hit 172 tokens. And what's going on there is that we are writing that to the cache and then as the conversation continues with the agent, great. we're just able to reuse those tokens and get that 90% savings. And so we've already kind of about half the cost of our agent because we've gotten that kind of about 58% cash hit rate. So nice, okay, this is a good first step to start out with. Let's take a look at maybe the rest of this dashboard. Wait, then there's five OKRs. I give you a million tokens of context and this is where we get, okay. Okay, so turns out we've filled up the context window and now we've got a little bit more work to do. So let's switch back to slides. What we need is context engineering. And what context engineering is, is the art and science of figuring out what context you expose to Clod to give your agent the best performance. And again, I said this earlier, I'm gonna repeat it many times. Look at your transcript to see it, to see what the models are actually seeing. And that will, I think, be incredibly illustrative to see. Is there a lot of stuff that you really don't need to be passing to Claude? Or is there a lot of relevant stuff that's keeping it on track? And we'll go through today. There's many techniques. We'll go through three primary ones about how you can implement better context engineering on your models. And so to start, the first is around figuring out what tools we pass to the model and kind of narrowing that down. The second is about narrowing down what results from those tools we share to the model. And finally is that conversation continues in the history. It's about keeping that conversation going for kind of almost unlimited feeling context. So let's go through these one by one. The first is tool search. Agents use a ton of tools. It's one of the magic of making great agents is tool calls. And cloud models are especially capable of leveraging those tools. We see tens, even hundreds of tools used in Agents today. Now, one of the problems, though, is when you define all those tools and you pass them to the model, what ends up happening is they fill up a lot of your context, and that leaves a lot less room for the actual work that needs to happen in the model. And so what you're looking at here is that, you know, that context fills up after just a couple turns. So instead, we have Tool Search Tool. And what tool search tool does is it figures out, you define all your tools up front, but we only pass the model a tool that tells it, okay, when you think you might need a tool, call this, and it gives then a list of the tools, and then only then does it put into context the actual definition of the tool. And you can see here on the bottom row, on the middle row, rather, we're just passing in those orange tools right when needed, leaving a lot more space for the actual stuff you need. And Lovable tried this. You heard from Bobbian earlier on this stage, and he was talking about how, he's told us with tool search that they actually reduced their overall token conception by 10% with just this solution. And importantly, it actually improved the performance of their model as well. That they saw that because you're putting less gunk into context, you're putting more relevant stuff there only. It turns out the models actually perform better. So they've rolled this out to all their users. Okay, so that's tool search. What happens about what you do with those results from the tools? That's the next one is programmatic tool calling. For programmatic tool calling, what the idea here is how do we curate that content that's coming back from tools just to what's most relevant? And the kind of, the insight here is it turns out Cloud's really good at writing code. If you've used Cloud code or any other coding tool that leverages the models, you have been able to experience this And we use that to your advantage here as well. With, we have Cloud just write a simple Python script that can actually call those same tools, get the same results, but instead do a little bit of work to curate that content and then send it to the model, just what's most relevant. And folks like Quora have used it. They've used it with HTML content where they've been able to strip away all the stuff that's just irrelevant, keep the part that's relevant, and seen the performance of their models improve. Okay, so the conversation continues. You've probably heard from folks on this stage before about how, from like Lisa and Jeremy, about how we're seeing the models able to do even hours of autonomous work today. And so inevitably you're gonna hit that million context threshold. And what with compaction is a way to allow you to continue that conversation instead of getting halted to a screeching stop when you hit that full context. What happens is it summarizes the context with your prompt, shifts it down to lower context, removes the turns that are no longer relevant, and then the conversation continues, rinse, repeat. And from there, you get this almost feeling of unlimited context, keeping the models on track through compaction. And Hex has used this, that since using this, they've been able to simplify down their code, and seeing that their model is able to continue to perform nicely. Nice, so let's go back to the demo. I think then we've, we've got a couple solutions of context engineering. Take us forward, let's see what we can do. So right here, you're seeing immediately that we've now seen that, see that context window bar kind of on the left hand panel? You see it's going up a lot slower because we're putting in less than the context in each turn. And what you'll notice is that when it hits around 400K, just right there, you notice how it went back down? What happened was that we hit a threshold, we've set it to 400K, I'll get into that in a bit, and then it compressed it down with compaction. So, hey, this is great. Let's take a look and see what we can do with, like, how's this working? Let's take a look at tool search, maybe? Can we just find one in there? Yeah, there we go. So we want to get our hero retention metrics. Remember, this is an analytical business. We need to understand retention. And so it calls the models and it asks them, OK, what are the tools that can help me get my hero retention metrics? Turns out there is a hero retention metrics tool. And now note in each of these, you can notice that that one is 14,000 tokens in the schema. The hero list is 6,100. The next one's 9,300. These are large tools. And that's totally normal for a lot of agents. And because we haven't put them into context, there's a lot less room that's being taken up by these tools that we won't need. And instead we call, okay, here are retention metrics, that's the correct one. If we go down to the transcript to find the, here we go, great. So what you see here now is these are the exact definition of that tool. Just that tool, none of the rest. And great, now the model can use the tool it needs and it leaves room for everything else. So hey, we've got the tool we need. Now what happens next when we get the results? Let's take a look at, yeah, right here. This is, so Gong is, if you're not familiar, it's a tool that kind of records your sales team's conversations. So you can look at the data, analyze it, see how the field is reacting to maybe the product you've launched. And these three-year-olds, I love them, but they really care about their green balloons a lot when we send superheroes to them. They go on and on for 30 minutes, even 60 minutes sometimes, And you're just not that relevant, frankly, when we're putting this dashboard together. I don't need all that data. I just kind of need the general sentiment, the vibes, of how they feel about the conversation. So with programmatic tool calling, the model has created a nice little script here, and it first looks at the first 2,500 tokens, characters. It's there, just understands what's the structure, and then it realizes all we want is the aggregate sentiment. And so it then writes a simple way to loop through, get the aggregate sentiment from the variety of calls and stream that in to the dashboard. And great, what previously was this massive result block that you just saw actually is now down to just the parts of it that we really need for what we need to do right now. Again, keeping that context narrow. And remember, this session's not about a cool demo. There's lots of cool demos in AI, love them too, but the kind of core thing here is about actually putting it into production. And in that case, you need to really make sure that the context really matches what's actually needed to make your product successful. So, okay, that's the second one. Let's talk about the third one, which is compaction. So we hit that 400K threshold, and now we need to take it down. We have chosen 400K. Now, I launched a million contexts. I'm a big fan of the million context. We know lots of scenarios where a million context is great. It might be for your scenario that the right combination of intelligence, cost, latency, is not a million. start with 500K, 400K is often a good starting point, but it changes by model. And what we've done is we set that threshold, and then we create the summary. And you get to create your own prompt to help guide it. And then it just puts together the key facts that's needed to summarize where things are, keep the conversation on track without losing the wrong context, and adds that in here. And then great, moves on to the conversation. So super, we've implemented context engineering. For those of you with your accounting eyes, you can see the cost is down to about 11 pounds. We've reduced it down about a third from where we were earlier. But I have to confess one thing. This has been a tough business. This one super hero you'll hear about, CryoThing. He does a lot of stuff that just increases my insurance premiums and it is a tough business. like margins are really, really thin. And yeah, we need to get the cost down. Every time I'm loading this right now, it's costing 11 pounds. That's pretty high. And what model are we using on this? OK, we're using for Opus 47, good model. But it's high intelligence and therefore also higher cost. I wonder, could we maybe figure out how to do this with Sonin and Hykubit? Do it with closer to Opus intelligence? Let's switch back to slides and see what we got. Okay, so we've got advisor strategy. Of course, I had a solution to that problem. So the idea behind advisor strategy is that you have the model, the agent is run with an executor that's Sonnet or Haiku. And the kind of insight here is that it can kind of figure out what to do with all these different shapes, except when it kind of, you'll see in a second, it encounters this sort of weird oddball shape, and it doesn't know what to do this, and it just calls the advisor, asks for what to do, and then it gives the advice on what to do. The insight here is if you've ever worked with development teams, you've probably noticed there are senior engineers paired with junior engineers that make that junior engineer so much more powerful that junior engineer is still the person hands on keyboard getting work done, but with the coaching, the mentorship, the code reviews, the help on architecture from the senior engineer, they were able to actually achieve so much more than they otherwise would have, sometimes even approaching what someone with that senior engineer skills could have done solo. And that's the kind of same idea here, it turns out that same logic works with models, that pairing the kind of sonnet haiku gets you kind of the prices of those models approaching even sometimes opus intelligence, a big intelligence boost. And folks like Bolt have used this, They've kind of gotten better architectural decisions where they can see on complex tasks, it improves performance. While on less complex tasks, it has no extra overhead. It just doesn't call the tool. And so it's a pretty nice trade-off. It's like a Pareto optimal way to improve your cost and intelligence performance. So let's go back to the demo machine and let's take a look. Let's see, let's go ahead and implement it. Okay, so we're starting at just about 11 pounds, and you'll see that bar is going up a lot less fast. Of course, we are using Sonnet. Now, as you can see, there's Sonnet 4.6 with Opus 4.7 Advisor. Okay, so it kind of seems obvious, of course, the cost is going to be lower. We just switched to a more inexpensive, I guess, higher value model. And the question is not, is it better quality, is it lower cost, is it also better quality, is it also losing intelligence on this dashboard? And so if we want to scroll through, I think that there was a, so we have this one contract that we've been struggling with, we really need to land it. I don't know if we're going to raise our next round if we don't get this. It's the Metropolis renewal. This has been our big customer from the start. They've been big supporters of HeroCorp. I need to keep them on board. And so, Sonnet looked through the transcripts, was like, I think the renewal is on track, we're doing well, but this is important, so let's ask Opus. And this is one of those ways that the advisor tool can be used. You can pass it a transcript, it'll take a look, anything wrong, and it'll report back of, hey, you might have missed a thing or looks good. In this case, Sonnet said it was green, but Opus looks through and it looks through it deeply, more fine-eyed and is able to say, actually, the mayor specifically wants cryothing. That guy was increasing my insurance premiums. Yeah, he specifically wants him. He's too good to lose, I know it. It's anyway. He's great, but he specifically wants him, but he's just unavailable that day. And so, actually, this is red. The renewal will not go through if he's not available on the day of the big event, they want at City Hall. And so this is a watermelon, if you've heard that term. It's green on the outside, but deep, deep red on the inside. And so what happens is Opus overrides Sonnet and lets us know and catches this far. So that's good. Okay, this is great. We were able to kind of recover that intelligence and have good performance. So okay, to backtrack here, what happened in this dashboard? We started at over 10 times the cost, and we brought that down through one prompt caching. We figured out what our prompt cache rate was, zero. We implemented it. And by the way, one great way to implement it, again, is the Cloud API skill within Cloud Code. It helps figure out a lot of that logic for you. We looked at the transcript then to see exactly what's going on, and we then implemented context engineering. First, reducing what tools are sent to the model. Two is then curating what data from those results gets sent back to the model. And then finally, helping create almost unlimited context with compaction. And then the last part was then helping reduce that cost further while preserving intelligence through Advisor Tool. So we've done all this. The Opus has given us a nice little button here in the air of AI agents, a CEO. I just click buttons nowadays. So should we save the metropolis contract? Yeah, let's do it, let's do it. Go ahead and click the button and metropolis is saved. We're gonna get our next round, I think, I hope. Let's see. If the VC's in the house, we can talk afterwards. So let's switch back to slides. Okay, so key takeaways. Again, to review, what did we cover here today? First thing, prompt caching. If you do nothing else, prompt caching. figure out what your prompt cache hit rate is, and implement it. Cloud API skill is a great way to get started. Then we did the context engineering pieces. We curated what's getting sent to the model, both from tools, from results of those tools, and then compacting it as that conversation grows. And finally, advisor strategy. A Pareto optimal way, in many cases, to get better cost and intelligence. A great trade-off. Now, I've talked about things that have launched literally in the last 24 hours here. So the cloud platform is evolving fast. I'm excited you've been here for this talk, but your learning journey is not over. We're gonna continue to try to do our best to make sure you can build not just great demos, but real production agents that work for you and your customers that let you build a business on this platform. And we're gonna do that continuously. That's our commitment to you. And that means keeping abreast of all the things that we're launching. Frankly, I didn't even have enough room on this slide to put everything that's launched in 2026. This is just a subset of what we've already launched this year. I'm particularly excited about automatic prompt caching, a one-line way to implement prompt caching, if you've never implemented it, or the cloud platform on AWS. This one's very cool. It's this whole platform that we've talked about, available where I know a lot of folks in this room use our models on AWS. It's all right there. So, yeah, this is a great way to get started. I'm excited to see what you guys build, and thank you for coming. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.