Token management is just part of the job now.
I've had to learn it the same way I learned to manage a budget, except the budget is compute. Which tasks I hand off to a smaller model. When high-thinking mode is worth the wait versus quick-think. What gets queued to run overnight instead of watched live. Where the 5-hour reset window sits, so I'm not tapped out by noon.
I hit the ceiling constantly. Mid-task, and the model tells me to come back in a few hours. So I drop to a smaller model and stretch what's left. Batch up questions ahead of time so I'm not wasting a turn on something I could've asked in one shot. If I'm rationing tokens at all, that's the tell — I'm using this enough for the cost to actually matter.
Right now we're in a subsidized window. The price tag on these tools has nothing to do with what they actually cost to run — it's a land-grab number. Same play Uber ran in 2014: cheap rides, every one of them subsidized by venture money, right up until the market was locked in and the prices stopped being cheap. Somebody's paying for all this AI infrastructure eventually. It's going to be us.
Then Jensen Huang says a $500,000 engineer who isn't burning through $250,000 a year in tokens should worry him. Half a salary, in tokens. Not as a nice-to-have — as the expectation.
That's the reframe. Your value as a knowledge worker starts tracking how much AI you can put to work, not how hard you personally grind. Throughput over effort.
Same math applies to a solo operator, just smaller numbers. The question isn't "can I afford this" anymore. It's "how small a version of me can still deploy the most tokens at the right problems."
Short term, I think it gets worse before it gets better — subsidies dry up, prices climb, the cheap era ends. Long term I'm betting the other direction: the compute build-out is real, local inference swallows the routine stuff, frontier models get reserved for what actually needs them, and per-token cost falls even as total usage explodes.
Hit the wall again this morning. Back in a few hours.