📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Why Claude Keeps Hitting Usage Limits (& how to fix it)

Eliot Prince26:12

Transcription

I've been getting really frustrated recently because I felt like I was hitting my clawed session usage limits quicker than I used to. And I assumed it was because I was just really hammering things like clawed co-work and working it really hard. But it's frustrating for Claude to just stop working halfway through a task cuz that kind of defeats the promise of AI of just not getting tired and and being able to work more efficiently.

But it turns out Claude have actually updated the amount of work that you can do in one single session with your Clawed account, whatever subscription you're on. So if you've been hitting your limits more frequently, then I want to go through all of this in this video. I'm going to show you the changes. I'm going to talk about tokens, context, and help you understand all the limits. And then I've actually put together something I've been working on a week: my 14 tips to solve this problem. 14 different things you can do to use Claude way more efficiently so you stop hitting and bombing out on your usage limits.

Now, what's been going on? Well, you can see this quiet post on Reddit here from Claude updates on session limits that if I zoom in a bit, you see, "To manage growing demand, we're adjusting the 5-hour session limits on the subscriptions during peak hours. Your limits remain unchanged, but during peak hours, aka 5:00 a.m. till 11:00 a.m. Pacific time, 1:00 p.m. to 7:00 p.m. GMT." So, in the afternoon for me here in the in the UK, more of the morning over if you're in the US, you'll move through your 5-hour session limit faster than you did before. Your weekly limit is staying the same, but essentially, you're not going to be able to do as much work during peak hours.

You can think of this almost like the public transport system or train tickets we have here in the UK. During peak travel hours, during commuter hours, tickets are more expensive because there's more demand. But at quieter hours, they're cheaper. And this is kind of what's been coming to AI and particularly Claude, especially with their massive growth in users. So, it's super, super frustrating if you're one of these people that's been effective. And I'm not going to sit here and just say, "well, just use AI or Claude at a quieter time of day" because that's not helpful. Because it's busy for a reason at those hours, at those peak hours, is because that's when we want to do our work.

So, let's break this all down, start to get you to understand how this all works and then give you the best advice to make your claude work more efficiently.

So the first thing you need to understand are usage limits and length limits. And these are two different things that are actually going to combine to understand how to work more efficiently. So if you go into your claude and you jump down into the bottom right corner here and go into your settings, you will see a tab on the right. You'll see usage and this is the place that you really want to monitor because you'll see your plan usage limits. What you have and this is the key one to understand is your current session and you can see here it resets in 49 minutes and I've used 677% of my allowance during this session.

Now sessions work in 5-hour blocks. So from the time you send your first message you then get 5 hours of this working session which has a caps limit. Claude don't actually give exact numbers on those limits, but we can start to dig a little deeper to understand how many tokens and how much context we're using.

Now, what hasn't changed is our weekly limit. You can see here I've used 62% of my weekly limit. It's Friday afternoon and it resets at Monday at 10:00 a.m. So, I've got plenty to play with through the rest of the day and through the weekend.

Now, at the bottom, we've got extra usage as well. Now, this is the first thing I want to talk about because extra usage is the way to keep Claude running if you hit your current session session usage. So, if you've ever been working in a chat and it suddenly goes, "hey, you've used up all your session. You're going to have to wait 3 hours until this time to continue working," you can get around this by just putting a little bit of credit on your account. So, I do have a little bit of credit on my account and I've spent an extra eight pound this month and that's often when I'm if I'm halfway through a video and I need to keep working like I don't want to have to wait three hours to finish shooting or if I've got it halfway through a task. It allows Claude to actually keep working on this extra usage for a few cents to actually finish that task it's working on rather than stopping and then just like being left half finished work which is the most frustrating thing.

So when we talk about usage limits, the main thing we're trying to understand here is your current session over a 5-hour block and how quickly you're eating up that session limit. That's the frustrating thing that we're talking about today.

Now the other thing to understand are length limits. These are also called context windows and those this is the amount of information that can be holding held in one chat window and Claude can process at once. Now, most plans have 200,000 tokens per context window or per chat. And every time you send a message or every response Claude gives, every file you attach, every job it does, that all eats part of your context window or the tokens that are essentially the currency of powering Claude.

Now, when I say 200,000 tokens, it's kind of hard to quantify, but tokens are you can kind of think of them as roughly words, but they don't exactly translate one token to one word. I think the rough estimate is one token is roughly about 3/4 to half a word. But I just want you to hold this tokenized concept in your mind. And as I said, every message you send and receive costs tokens. We have 200,000 available in most of our contact windows. Sometimes this can be pushed to a million per chat, which is why again we're going to delve into all the settings and make sure we're we've got the right options on. And what's also important is not all of Claude uses the same amount of tokens. If you're running some of the higher-end models, Opus 4.6 for example, that's going to cost five times more tokens than Sonnet.

Now, the first thing I want to show you is if we jump into a new chat here in Claude, the first thing to understand is what model you're using. So on the right hand side here, you can see Opus 4.6 is going to burn tokens five times faster. So you're going to use up your usage in a 5-hour window five times faster. Now I usually just switch down onto Sonic 4.6 which is the most efficient for everyday tasks.

Now you might think, "Well, I'll always just use the cheapest one," but you should also understand that for some more complex tasks or some more intensive tasks, it it's not always the most efficient. You need to think about imagine uh moving people in rush hour. For example, you could use haiku to move a 100 people from one train station to another, but you're doing them individually one by one in small trains or on motorbikes or something. But actually, you'll notice that when you're moving large amounts of people, hauling a lot of stuff around, it's actually more efficient to do it in one big batch in one big train carriage or one big truck to do it all at once rather than making lots of journeys back and forth doing things one by one. So, sometimes you might actually find that just using Haiku for everything is taking a lot longer. And actually, it's not actually that much more efficient when you break up break down how many trips back and forth and things it's having to work out. So, Sonnet's usually the best and I will push it up to Opus if it's something particularly intensive I want to do and I've got plenty of usage ready to go. So, that's the first thing to understand.

The second big tip here is to go forward slash and then you want to type in context. And running this context command is going to open up a secret portal in Claude here to show you exactly what's going on and what's using your tokens under the hood. Look at this. We've got all of this information and this is actually showing us these plugins are using tokens. These different tools and skills are using tokens. We've got all these MCPS and connectors that are using stuff up. See Apify here is using a ton of tokens. And this is what's called your context usage. But again, it gives a breakdown of your tokens. So in every chat or in this chat, I have 200 tokens available to use. Once I use up 200 tokens, the chat isn't really going to be able to reference that previous context. It might do this thing where it compacts and tries to actually use some autoco compact buffer here to compact all the information I've given it and chatted through so far, but it kind of loses a lot of the detail and just pulls out the key facts to carry on working in our chat. So, it won't be as well informed if you start using Claude past these 200k tokens.

But what we can also read into this is actually the tokens that are burning up our usage allowance in any 5-hour window. So you can see my system prompt. So my custom instructions on top are using 20,000 tokens straight off the bat when I even just load up a chat. I've got system tools. These aren't using very much. Not worth worrying about. got MCP tools which are taking up 45,000 tokens potentially although they do defer some of this limit so it's not the full whack but you can see keeping on all my MCPS and all my tools might be actually burning credit every time I'm running a chat even if I'm not using them and then we've got other things in here memory files okay that's fine that's not too much we've got clawed skills I've got a lot of claude skills running that are using up a little bit of token so you can start to see things down here that actually we might want to start tidying up. And we're going to go through more about this in our video.

Now, what I would do here is simply ask like what is burning the most tokens in my chats to see like, is there anything in here that's burning this context and these tokens every time I run a prompt? And it starts to give you some insight into what's actually going to be burning all of your usage. So, if you're finding your usage limits are running up pretty quickly, there might be some ways you're using this, some things you've got switched on that are causing this problem.

One of the biggest factors and the next tip is toggle off thinking. So in a claude window for example, you have all your little plus button here, your tools, your connectors and other in-built tools included. So things like web search and research and these extra tools you're layering on top. If you're not using them, they might be using up some context as well. But the real killer is this extended thinking button. Some people turn this on and they forget about it and they just let Claude just go to town and have extended thinking on all the time, but it's actually burning a lot of tokens cuz it's going through extra steps. It's putting extra outputs. It's computing more. You're probably working for back and forth with it a lot more. So, if you've got extended turned on, make sure you turn extended off in your clawed chat to save tokens.

So, start by looking at those basic settings. Try to maybe stay in Sonic more often. switch off your extended thinking. Run that forward slashcontext to see what's actually eating most of your context and tokens through your usage.

Now, we can level things up a bit more and actually talk about your conversation habits. As we just discussed in your chat, you have all of this stuff going on that runs straight out the box. I've already run 30,000 tokens of my 200,000 chat allowance for this task. But the longer my conversation gets, all of that previous chat history is an overhead in every prompt I run. So, every time I run a prompt, it's going to remember all of the stuff that went before, and that's going to add to that token usage overhead and drag things down even more.

So, one great tip is to start a new chat per task, particularly if you've gone from different ideas or you're working on slightly different things. So, if you go from analyzing spreadsheets to write writing marketing content, doing that in the same chat is going to keep burning your usage limits more quickly because you've got that context accumulating with every message in the thread. It's like carrying extra baggage with you. So, by the time you get 20 or 30 messages in, not only is your thread and your chat just a bit of a mess, you're burning significantly more tokens. So, start one task or one chat for particular tasks that have the same context and same information behind them to reduce that extra overhead.

The second thing here is to be specific with your prompts. As you've just seen, starting a task with anything is going to burn a significant amount of tokens anyway. There's no way around that. But if you're just kicking things off with things saying things like with tell me about this document and you have Claude go through a full massive PDF page of information because you want a summary, but you're actually looking for something very specific. Tweak your prompt to be more specific about what you want it to do from the off. Don't get it to summarize a whole document and then zero down. Say like, actually, no, I want to summarize the financial risks in section three of this PDF so Claude knows exactly where to look and where to get going. That spec specificity is efficiency.

Now you can level this up as well. As we've seen, every prompt is burning some token overhead with all the tools and context running. So batch your requests into fewer prompts. So instead of going one by one and giving it one fix at a time, can you fix the typo in paragraph 2? Also then shorten the intro. Add a CT at the end. Try tightening it up and giving it three things. Fix the typo. shorten the intro sentence and add a CTN CTA at the end of the link to the vault. That way, you're grouping your message. It can go off and do all the work in one shot rather than three messages with three sets of token overheads and burning up more of that context window which is going to tap out at 200k and then the quality of your conversation is going to significantly drop.

Then, as we just saw, we need to start killing some of these background tools and connectors. This is the one that surprises people the most because you've got things like extended thinking, you've got web search, but then you've got your connectors too. For me, if I go into my plus button here and start to see my connectors, you can see I've got things in here that are just on all the time. I've got my drive, search, base row, firecrawl, gmail, notion, apply, claudin, chrome, control, chrome. This is something I spotted and this was a massive thing for me. Claude in Chrome running. But previous to installing Claude in Chrome, I had a custom control Chrome MCP, which don't worry about that if you don't understand what that means. But essentially, I've got two connectors or tools that are both trying to run Claude through Chrome. I don't need both of them, but they're both eating my token usage. So, I can just turn one off. Well, actually turn the old one off and keep the new one on and immediately start saving tokens and context through my usage limits.

Now, these are all slightly small things. They're not going to be game-changing, but once we stack all of these 14 different things up, you're going to see massive improvements.

So, step two, after you've gone through your settings, start your usage habits or your chat habits of working on a very particular task per chat, being very specific in your prompts, batching your requests rather than going one by one by one, and then disabling tools that you don't actually need all the time.

Then we can move on to tier three of a bit more of the setup behind the scenes or how you set up your Clawude workspace, either Claude chat or Claude co-work depending on where you prefer. Now, this is one of the most underrated things going on here. I'm guilty of this. Everyone's guilty of this is just throwing documents like PDFs at Claude and just uploading PDFs or pointing co-worker a folder full of PDFs and just saying there's all the work. Now the problem with that is that is quite heavy work particularly if you're using them a lot and doing heavy research. And for a lot of people if you're throwing a lot of PDFs it could be up to 80% of a session just kind of extracting information out of out of PDFs. And I've seen a lot of people who feel like their whole session evaporates in one go because they've given it a ton of files and then it's like, "Okay, I've read them all. Now I'm out of time."

To fix this, if you know you're going to be dropping a lot of PDFs and heavy documents in, you can actually spend time converting them into a markdown or plain text. So if I was to go over to something like chat GPT or no, in fact, because I don't really use chat GPT anymore, what I would do is I would go to perplexity. I have perplexity running and I can use some of my usage in here rather than burning it on claude where I actually want to do the work. So to convert these files I could even go and run Claude Sonet in here but I can leave it on um the best selector model but perplexity has access to GPT Gemini Claude all of them in there. Then I could give it a PDF or a set of PDFs and just ask it to can you convert this PDF into markdown so it's in a lighter format. And I'm going to use perplexity here to actually extract the information and put it in either markdown or a text file. Markdown is the best. You can see here it's kind of the love language of LLMs that they prefer reading. It's easier to extract all the information. Just gives me all the core stuff out of this PDF rather than any of the design, the images, any other heavy stuff in there. I can just take this information and take it into Claude to actually start working with. I could paste it in there. I could get Perplexi to do this over a number of files and actually get them as downloadable so I can put them in a space for co-work to work with as well. However you want to do it, but getting your data in a more readable format is well worth doing before you actually start working in co-work. If you want to use another platform to do that, that's going to help you not use usage limits up.

If you're intent on doing the work now in Claude, then you can stack this and use Claude projects if you're using Claude chat in particular because it uses rag retrieval augmented generation and that means it's only pulling the relevant content into your context window rather than using every file. So before we might have been just dumping a load of PDFs in there to start working within the normal chat. Now we can actually process those files into a more readable format and we could put them into projects. So we fill our projects with markdown files rather than he heavy PDFs as well. We're getting a fundamental shift in how we're using Claude. We're not using all that context on heavy files immediately. We're actually just picking the core readable information as we need to use it.

And that's helpful in projects as well because you'll see if we go into something like Claude projects, you'll see one, we have text files here that are just easily digestible by Claude. Two, we can write custom instructions about how we want it to act every single time. So we're not wasting context, credit, tokens, whatever. Actually reexplaining ourselves, setting it up, telling it what we want to do. We can almost write these so it's just like here's what I want to achieve and it knows what to run because it's got all the rules ready to go. Three, you can have memory running in your project as well. So you can just pick up where you left off and not forget where you were at. And you can install all of this sort of stuff, memory and instructions. You can actually make that work in Claude Co-work as well. You can get co-work to build you um instruction files here. You can see I've got Elliot Prince's co-work. You can get it to set up memory files, everything like that as you work.

Then what happens is a lot of people actually end up writing really long custom instructions or project instructions that can be thousands of words long. Particularly if you're using Claude code or Claude Co-work, you can get way over the top in adding things into your instructions. Actually, you want to keep your custom instructions less than 500 words or as short as possible and to the point as possible because every single message you're loading and reading your custom instructions. Every time you load up a new task, you're loading your custom instructions, which again is adding to that overhead. So, take the time to check your custom instructions wherever you've got them to try and trim them down and keep them as efficient as possible. Again, basically, the less words we can have clawed processing, the less tokens and usage we're going to use.

Then of course, remove any pointless files. If you've got files that are sitting in your co-work folders or your Claude projects, just get rid of them if they're not needed because it's just a waste. Claude might not actively be checking these this context every time, but it is just more clutter for it to filter through and actually work out what's going on. It might accidentally start going through stuff that isn't relevant.

Once you've done all that, then I would suggest building Claude skills. This is for a couple of reasons. Now the beauty of skills is if you do the same type of task regularly, same stepby-step thing or you need the same output every time, whether that's for me it could be like creating invoices or generating reports or writing introductions in a particular way or publishing things here and there. Skills have the exact step-by-step process baked into them. So you don't need to reexplain every time to Claude or you don't let need to let Claude go and figure out how to do something each time and be like, "Okay, I think I can do it like this. Actually, that wasn't quite right. Let me re-angle and try and do a different thing or spend context token working out how to do stuff for you. This just hands it the recipe on the plate and it can go and just execute the same t same way every time."

Now, the clever bit with skills are they only load the description into your context. So they're not burning tokens actually reading the whole skill every single time. So when you go into your customize area of Claude here and you go to skills, you'll see here for example that this teaching dashboard that I build, I've got this built as a Claude skill. So I can give it the information and I just say, "Go and build the dashboard." Now the only bit of this skill that actually loads off the top is the description. So, one, keep your descriptions quite short, but two, it just gives Claude enough context to be like, "What does this thing do and when should it activate?" And that's all it needs to know in its context. The rest of it here, it only loads the rest of this information or these assets and files in here when I actively only need it. It's not going to load all of these skills all the time. So if we go back to our original context, yes, I've got 20 or 30 skills running probably, but actually it's only using a little slice of tokens to actually run. So they're super useful. Highly recommend start building Claude skills to help you work more efficiently rather than hit these context limits all the time.

And then once you've got all that done, we can then move on to tier four, which are a couple of sneaky tricks I like as well. Now obviously the first one was shift heavy work off peak. So for me the peak hours where remember Claude are essentially tightening usage restrictions or each token doesn't go as far and does as much work. 1:00 p.m. to 7:00 p.m. in the UK is when those restricted hours are. Those are the peak hours. So I should be planning and it's 1:43 now. So I've done most of my work in the morning. Um, I should be planning to do all of my heavy clawed work first thing in the morning to take advantage of the the cheaper rates essentially. It's a bit frustrating, but there's nothing I can do about it. That's the world we live in. So, schedule your work if you can to actually or particularly if you've got scheduled tasks include, schedule them to run outside of these peak hours.

Number two, this is what I call the session reset trick. Now, your 5-hour window starts the moment you send your first message. Not necessarily when you log in or not on Anthropic and Claude's own top on clock or when you open a chat. It's actually when you prompt Claude. So, you could be quite clever and send a throwaway prompt earlier in the day or a few hours before you actually want to start work and just say, "Hey, let's get started." That starts your 5-hour window starting. Then you can come back a couple of hours later or maybe an hour before your context window or your usage window is going to expire and then start your work because what you'll find is you'll be able to do a fair bit of work and then by the time you actually get towards your usage session if it's something particularly heavy it might be about to reset or you'll have got your fresh allocation. So, you can almost get two sittings you can get 10 hours worth of usage if you spam those two times. So, that's a neat little trick to just a small thing, but actually get you around that session limit problem almost immediately to double what you can do in one sitting.

So, those are my 14 tips to help you reduce your usage. Remember the difference between usage limits of like how long you've got in a five hour period versus looking at your context window to see actually how much time have I got left or how many tokens have I got left in this particular task or chat before it's going to start to decay and use the forward slash context to see what's eating up your tokens.

Then you can make sure to maybe try and switch to sonet more often than running an opus. Toggle off any extra tools like extended thinking. Try to batch your tasks into things that are in the same vein or require the same information in the same context rather than being scattered and jumping to different things in the same chat window. Be very specific with your prompts. Batch your requests and your prompts into multiple requests in one prompt to stop that extra context and token overhead. And disable any idle tools and MCPs and connectors that just you don't use or you don't use regularly. So, they're not using up your limits.

Then start working a bit smarter. Process documents and get things into a markdown or a text format before you start working with them. Use projects and keep them clean and efficient with nice clean short set of instructions that aren't overwhelming. Keep them clean without any idle files or things you just aren't relevant in there and start building out Claude skills for your repeatable workflows so that Claude can just execute an exact process every time rather than going off in different directions. Then shift your heavy work or particularly your scheduled tasks for outside peak windows and use that session reset trick to take advantage across two session windows.