What’s Actually Eating Your Claude Usage (Hint: It’s Your Context Window)

by | Jun 16, 2026

By Loren Bartley, Impactiv8

Have you ever found yourself in a situation where you’ve had a single Claude chat open all day, and you’ve just kept piling things into it? Draft this, now sort that, now have a look at this other thing. By late morning, it’s slowed right down and started forgetting things you’d told it an hour earlier. Your first thought might be that “Claude’s having an off day.”

I’d hate to break it to you, but it’s not Claude. It’s you. When you let a single conversation balloon into a monster, that monster quietly eats away at your usage, dragging replies out, and making the work worse, all at the same time.

In my last post, I showed you how Claude’s usage limits really work and the free timing trick I use every day. This is the other half of that story. Once you know how to time your day, the next question is what’s actually draining your budget in the first place. Nine times out of ten, the answer is your context window. So let me explain what that is, why long chats cost you more than you’d think, and the simple habits that make your budget, your speed and your quality all better at once. Best part: it’s free.

Two limits, and people muddle them

Quick bit of plumbing first, because this trips almost everyone up.

Claude has two different limits. Your usage limit is your budget over time, how much you can do before you wait for a reset. That’s the one I covered last time.

Your context window is how much a single conversation can hold in its working memory at once. They’re separate, but importantly, Claude Conthey feed each other. A bloated context window makes every message in that chat more expensive, which drains your usage limit faster. So managing your context window isn’t just about avoiding a “this chat is full” wall. It’s one of the best ways to make your usage stretch.

What a context window actually is (in plain English)

Picture a box. Everything in one conversation goes in it… Your messages, Claude’s replies, any files you’ve attached, your instructions, and the answer Claude is writing. If it fits in the box, Claude can see it. If it spills out, Claude can’t.

The box is measured in tokens. A token is just a chunk of text, roughly three-quarters of a word. To give you a sense of scale, most chats hold around 200,000 tokens, which is something like 500 pages. The newest models stretch to 500,000 tokens when you’re chatting, and in Claude Code and Cowork it can go up to a million. That’s a big box. So why do you ever run into trouble? Because of how the box gets filled.

You never start with an empty box

It might shock you to hear this, but even a brand-new chat isn’t empty. Before you’ve typed a single word, the box already holds the system instructions that tell Claude how to behave, any custom instructions you’ve saved, your connected tools, your memory, and the names of every Skill you’ve got available.

So you’re never filling an empty box. You’re topping up one that’s already part full. Once that clicks, everything else makes sense, including why switching off tools you’re not using actually helps.

Claude Context Window

Why long chats get expensive

This is the part that genuinely changes how you work, so stick with me.

Every single time you send a message, Claude re-reads the entire conversation to reply. Not just your latest message. The whole thread, from the very top, every time. So a chat that’s fifty messages deep costs far more per reply than a fresh one because Claude is hauling the entire history along with each new answer.

That’s the real reason behind the old advice to “start a new chat.” A fresh chat is a near-empty box, which is cheaper, faster and sharper. A sprawling one is an overstuffed box you’re paying to re-read at every turn.

It’s not just your budget. It’s quality and speed too.

Cost is the obvious one, but a stuffed context window costs you in two other ways as well.

The first is quality. As the box fills up, Claude’s answers actually get worse. It starts forgetting things you mentioned earlier and its reasoning slips. This isn’t Claude being grumpy, it’s well documented across every major AI model. They all degrade as the context gets longer. There’s even a name for it… context rot. Yuck!

The second is speed. Because the whole context is re-read on every message, a bigger box means more to wade through each time, so your replies come back slower.

Put those together and the lesson is lovely and simple… A lean, focused chat is cheaper, smarter and faster, all at once. You’re not trading one off against another.

Where Skills fit in

A quick word on Skills because they cut both ways.

The good news is Skills are clever about context. Only a Skill’s name and one-line description sit in your box by default, which is tiny. The full instructions load only when you actually use it. So a Skill keeps a reusable workflow tucked away until the moment you need it, instead of you pasting the same wall of instructions into every chat and bloating the box every time. Used well, Skills save context.

The watch-out is small but real. Having lots of Skills available nudges your starting point up a touch, and when a heavy Skill does fire, its full instructions plus any files it reads land in the box and can fill it quickly. So keep the ones you actually use, and don’t be surprised when a big Skill takes a decent chunk of the window.

How to tell when your window’s getting full

This is the question I get asked most, and the answer depends.

In regular Claude chat, there’s no neat little meter showing your context. The usage bars under Settings then Usage track your usage limit, which is a different thing. So for the context window itself, you’re reading the signs, like a warning that you’re near the length limit and should start a new conversation, and the tells I described earlier. If replies start slowing down or Claude forgets something you said a while back, that’s your cue. The box is getting full.

Claude Code is the one place with an exact readout. The /context command shows you precisely how full your window is and what’s taking up the room. But that doesn’t work in Cowork. Instead, you can simply ask Claude how much of the context window you’ve used, and it’ll tell you.

The practical rule for most of us is to not wait for a number. When the replies slow, or it gets forgetful, start fresh. Over time, you should start getting a feel for when might be a good time to say goodbye to a chat based on how much work you’ve already got it to do in advance of any degradation in performance. That’s the ideal.

Context windows Claude Chat, Cowork and Code

The habits that keep your box lean

Here are the moves that make the biggest difference, roughly in order of impact.

  • Start a fresh chat for a fresh task. This is the big one. New task, new chat. Don’t pile a brand-new job onto yesterday’s sprawling conversation.
  • Use Projects, and prefer a working folder over attachments. A working folder lets Claude pull in only the bits it needs. An attachment loads in full whether it’s used or not.
  • Keep your instructions tight and clear out files you’re not using. Lean instructions and fewer stray files mean more room for the actual work.
  • Switch off tools and connectors you don’t need for that conversation. They take up space, and they can confuse Claude when several do similar jobs.
  • Build a Skill instead of re-pasting instructions you use again and again.
  • Batch your questions into one good message rather than ten little back-and-forths.
  • Match the model and effort to the job, and turn off extended thinking for routine work.

Chat, Cowork and Code: same idea, different controls

The concept is identical everywhere. What changes is the size of the box and how you manage it.

In regular chat, Claude handles most of it for you, quietly summarising older messages as you near the limit. You steer it by starting new chats and using Projects. In Cowork, the box fills up much faster because every file Claude reads, every tool it calls and every Skill it runs adds to it. So a fresh session for each distinct job matters even more. In Claude Code, you get the biggest box plus hands-on controls to clear or compact the conversation yourself.

But the one thing that works across all three environments, is starting a new chat to clean the slate.

The bigger point

If you take one thing from this, let it be this… Your context window isn’t something to fear or fuss over. It’s just a box, and best practice is keeping it lean. Do that, and you get the trifecta:

  • A budget that stretches further
  • Answers that stay sharp
  • Replies that come back quickly

Pair this with the timing trick from my usage-limits post, and you’ll get the most value out of Claude.

And if you’re getting set up or weighing up a plan, you can get started with Claude here.

FAQ

What’s the difference between Claude’s usage limit and its context window?

Your usage limit is your budget over time, how much you can do before it resets. Your context window is how much a single conversation can hold at once. The limit that pauses you mid-task is almost always your usage limit. The context window is what makes a long chat slow and forgetful, and a fresh chat resets it.

How do I check how much context I’ve used in a chat?

In regular Claude chat there’s no exact meter, so watch for the warning that you’re near the length limit, and the tells, like slower replies and Claude forgetting earlier details. In Claude Code, the /context command shows you exactly how full your window is. In Cowork, you can just ask Claude how much of the context window you’ve used.

Does starting a new chat really save usage?

Yes. Claude re-reads the whole conversation every time you send a message, so a long chat costs more per reply than a short one. Starting fresh for a new task gives Claude a near-empty context window, which is cheaper, faster, and produces better answers.

How big is Claude’s context window?

On most paid chats it’s around 200,000 tokens, roughly 500 pages. The newest models support up to 500,000 tokens when chatting, and in Claude Code and Cowork it can reach a million. A token is about three-quarters of a word.

Do Skills use up my context window?

Only a little by default. A Skill’s name and short description sit in your context, so Claude knows it’s there, but the full instructions load only when the Skill is actually used. They’re a context saver compared with pasting the same instructions into every chat, though a heavy Skill that reads lots of files will take up more room while it runs.

READY TO GET BETTER RESULTS FROM YOUR MARKETING LEVERAGING AI?

Whether you want to master AI-enhanced digital marketing at your own pace, get hands-on with AI in a workshop, or have personalised guidance to fast-track your results, we’ve got the training, support and accountability to help you get there.

LEARN TO
LEVERAGE AI

A hands-on AI workshop designed to help you use AI to simplify your marketing, save time and get better results, with practical, step-by-step guidance, ready-to-use tools and AI skills you can put into action straight away.

MASTER YOUR MARKETING

Join the AI Mega Success Academy for step-by-step training, live coaching and a supportive community of business owners, so you can confidently grow your business through AI-enhanced digital marketing.

TAILOR YOUR
TRAINING

Whether you need one-on-one coaching or want a bespoke workshop to upskill your team, get personalised AI-enhanced digital marketing support designed around your specific goals, challenges and preferred tools.