Close the Chat on Purpose
The most useful habit in my AI practice is not a prompt. It is knowing when to stop the conversation, not the model.
I wrote a while back about the week I could not bring myself to open a new chat. The thread was long and warm and full of everything we had figured out together, and starting over felt like walking away from work I would never get back. So I did not. I kept feeding the same conversation, and it kept getting quietly worse.
That week did not happen in my first month with these tools. By then I had been using them heavily for years, every day, for real work. I understood the problem in the abstract. What I did not have yet was a system, and understanding a problem without a system for it is how you end up doing the thing anyway.
There is a moment in my own logs I keep going back to. The thread had gotten long, so I typed the obvious fix: can you compact the conversation please? The answer came back a few turns later, and something was missing from it. I told the model exactly that. You lost it on compact. Here they are. And I pasted back in the very things I had trusted the conversation to hold.
I did not read that as a warning at the time. I do now. That was the container leaking, and I was the one who noticed only after something fell out.
Because for longer than I would like to admit, I treated a context window like a container. You pour work into it, you are fine until it is full, and then it spills, so you reset at the top and not before. That is not how it works. A context window is a budget, and the budget thins the whole time you are spending it. The model does not hold full quality right up to the edge and then fall off a wall. It fades gradually, across the whole range, the way attention fades in a room that keeps getting louder.
I read the research on this not long after it came out, about a year ago now, and it said what I had been living, only more precisely. Run a frontier model across a growing pile of context and its accuracy slides down the whole way, a slope rather than a cliff. Quality can be measurably worse long before the window is anywhere near "full," sometimes at a fraction of what the label says it can hold. The reliable working range is a good deal smaller than the advertised one.
There is a second effect, stranger than the first. When you bury something important in the middle of a long conversation, that is exactly where it is most likely to get dropped. Models keep the beginning and the end and lose the middle, regardless of how much the middle mattered. The longer the thread, the deeper that middle gets, and the more of your own instructions sink into it. Which is precisely what happened the day I asked it to compact and had to hand my own material back.
So the long, warm chat I refused to close was not a container filling up. It was a position quietly going against me.
Once you say it that way, the move is obvious, because it is a move I already know.
A decaying position is the most ordinary thing in markets. You are in something, it is slowly bleeding, and it is not dramatic enough to force a decision. The amateur averages down. He adds to it, sits with it, tells himself it will come back if he just holds a little longer. The operator flattens it. He takes the small, deliberate loss now instead of the large, involuntary one later.
Keeping a fading chat open is averaging down. You feel like you are protecting the work. You are actually paying, a little at a time, in answers that are subtly wrong and mistakes you do not catch, until one of them costs you something real. Starting a fresh chat is flattening the position. It feels like giving up the thread when it is really just refusing to let a degrading one bleed you.
Knowing all of that did not hand me the routine. Between understanding the problem and having a system for it, there was a long stretch of fumbling that nobody puts in their posts about AI workflows.
For a while my fix was brute force. I opened every chat by dumping as much as I could about my situation and what I was trying to do, then told the model, item by item, to remember the things I knew were important. It worked, sort of, and it was mentally taxing in a way that added up. I watched people who write code for a living run their sessions and took the pieces that transferred. The routine below took me the better part of a year to converge on, and none of it is clever. That is mostly why I trust it.
I clear after every task, or every cluster of tasks that belong together. Not when the window is full, but after the job. If I have finished one thing and I am starting something unrelated, I take the clean slate, because a tight fresh context beats a long one dragging everything that came before it.
I do not wait for the top. If a job runs long, I reset somewhere around the point where a third of the budget is gone, because that is roughly where the slope starts to bite. The rule is the end of the task or a third of the budget, whichever comes first. Earlier than feels necessary. That earliness is most of the discipline.
I put what matters at the edges. The constraints I cannot afford to have forgotten go at the very start and the very end of what I hand over, never buried in the middle where the model is most likely to lose them.
And none of that means the work evaporates when I close the chat, because I stopped trusting the chat to hold it in the first place. My real habit runs the other way. I open a session by telling the model to recall everything we have discussed about a given project before I spend a single token on it, and I end a piece of work by asking it to summarize the thing so I can pop it into a new chat and keep going clean. The record lives in notes and files and the system that outlives any one conversation. The chat was never the asset. It is scaffolding, and you take scaffolding down once the building stands.
The last piece was getting the guesswork out. A meter sits at the bottom of my screen while I work and shows me exactly how full the context is. A small bar runs green while there is room, turns yellow the moment I cross into the reset zone, and goes red once I have pushed past it. I do not have to feel my way toward the edge anymore. I can watch the budget drain and close the chat on purpose, before it costs me anything.
The same meter, three states. The number on the far right is the running spend, which says the quiet part out loud: a long, tired context is not just lower quality, it is more expensive per unit of work. The bar fills, the answers get softer, and the bill goes up at the same time.
I will say one more thing, because the tooling around this changes weekly and it is easy to feel permanently behind. New models, new features, new things you are apparently supposed to master, and a low-grade FOMO humming underneath all of it.
I do not run my practice on FOMO. I run it on friction. The questions are always the same. What am I trying to achieve, and is anything actually bogging me down. If something is, I go find out how to eliminate it, and that search is where every habit in this piece came from. If nothing is, I keep going with what works, and I do not much care whether I look like a tech bro super coder while doing it. There is a middle ground between burying your head in the sand and trying every new thing the day it ships. Friction is how I find it.
The thing I got wrong that week was not a setting. It was a belief about where the work lived.
I thought the work lived in the chat, that closing it meant losing it. It does not. Anything that matters gets written down, into notes, into files, into the system that survives the conversation. I used to keep chats open the way I once kept losing trades open, out of an attachment that dressed itself up as patience. Now I close them the way I learned to close a position that had stopped working. Cleanly, early, and without the story.
Knowing when to start over is not abandoning the thread. It is the same skill as almost everything else worth having. You take the small loss on purpose, so you never have to take the big one by surprise.

