The Silent Tax of Long Conversations
TL;DR: Long AI chats slow down and start forgetting because every prompt re-sends the whole history. Where the thresholds are, and when to start fresh.
I was building a complex YAML configuration yesterday when I hit a wall. Not a syntax error, not a permission issue, but a slow, creeping sluggishness. The model was taking longer and longer to respond to each prompt. I tried to push through, pasting in failed outputs and asking for corrections, but the chat was becoming unusable.
It wasn’t until I asked the model what was happening that I realized I was fighting the architecture itself.
Every time you submit a prompt in a long conversation, the AI doesn’t just process your new input. It resubmits your entire chat history and re-processes everything from the beginning to establish context. As the chat grows, so does the computational load. Eventually, the model slows down to barely usable speeds.
This is the silent tax of context windows.
Two terms, since I’ll use them throughout: a token is the unit the model reads and writes in — roughly three-quarters of a word. The context window is the total number of tokens the model can hold at once: your prompts, its answers, and anything you pasted in, all added together.
The Thresholds
I’ve been tracking token usage closely, and the degradation isn’t linear. It’s stepped.
When I asked the model directly about its performance characteristics, the breakdown was clear:
- 0–100K tokens: Optimal performance. No degradation.
- 100–150K tokens: Still excellent, but it might miss minor details in the early conversation.
- 150–180K tokens: Noticeable context compression. It may forget specifics from the beginning of the chat.
- 180K+ tokens: Significant risk of losing important details from the start.
The model recommended starting fresh after hitting ~100K tokens, or before any critical system changes. I’ve found that if I wait until 150K, I’m already in the danger zone for forgetting earlier decisions.
The “Kickstart” Strategy
The solution is simple, but counter-intuitive to how we usually work. Instead of pushing through the long chat, you need to break it.
When I detect sluggishness or hit the 100K mark, I ask the model for a “Kickstart.” This is a summary of the long chat we just had, containing all the facts, decisions, and context a new chat would need to know.
I then start a brand new chat, paste that Kickstart as the seed, and continue working.
This resets the token usage to zero. It eliminates the dozens of failed versions the model was trying to process on each prompt. It also restores optimal performance.
Why This Matters
I used to think the context window was just a hard limit—a ceiling you hit and then stop. But it’s more like a pressure cooker. The longer you stay in the chat, the more the model has to carry with it.
If your AI is forgetting instructions or details, don’t just ask it to “remember.” Ask it to summarize the current state, then start fresh.
It’s not about the model’s capacity. It’s about your workflow.
Takeaways
- Watch the 100K mark. This is where performance starts to dip.
- Ask for a Kickstart. When the chat gets long, ask for a summary to paste into a new chat.
- Reset often. Don’t wait for the model to fail. Reset before you lose context.
- Use projects. If you’re using Claude, projects limit the context to specific files, which helps manage the load.
I’ve started doing this proactively. It feels like cheating at first, but it’s actually just respecting the mechanics of the tool.
The model is a builder, not a historian. Give it a fresh page, and it’ll build you something better.
I’ve been saying this for a while. This piece was pulled together by AI from 9 of my social media posts (07/2025 - 08/2026) and polished for Cairoglyphics.ai.