Four companies shipped a flagship model in a single week at the start of September. Meta, Google, OpenAI, and Anthropic, one after another. If you use any of these tools for real work, nothing on your screen changed that week, and nothing you had already built stopped running.
The pressure to move showed up anyway, and moving has a price.
What model fatigue actually is
The industry has started using a term about itself. CNBC reported last week that the median interval between major model releases has compressed from 37.5 days in 2023 to about 11 days so far in 2026. For OpenAI alone, the gap between launches went from 170.5 days to 49.
Eleven days is the number that matters. It means that if you set out to properly evaluate every significant model that ships, you would be starting the next evaluation before you finished the last one. There is no version of that schedule where you catch up.
This is landing on people who are paid to keep up, and it is landing hard. In July, more than a thousand employees across the major labs signed a petition asking for a more measured pace. Engineering teams are re-running benchmark suites every few weeks because the leaderboard has moved again by the time the previous run finishes. Some companies have stopped trying to evaluate everything and now cap their shortlist at about five models.
Those are professionals with budgets and staff. You are one person with a business to run.
The hours
A release every eleven days is roughly thirty-three significant models a year. Give each one a single honest afternoon on your own real work, call it three hours, and that is ninety-nine hours. Two and a half working weeks a year auditioning software instead of using it.
An afternoon is also not enough to learn anything real. The failures that matter do not appear in a quick test. They appear forty steps into a long job, once small errors have had room to compound and the thing goes quietly off the rails on task nine of ten. Finding that takes days of ordinary use.
The money
The hours are the hidden cost. The subscriptions are the visible one, and they accumulate faster than people expect because every individual number sounds reasonable.
Twenty dollars a month is the entry price at the tools most people compare. Claude Pro is twenty dollars billed monthly. ChatGPT Plus is twenty. Google and Meta each sell a comparable tier. Any one of those is an easy yes on its own.
Run three side by side so you can compare them fairly on your own work, and you are at sixty dollars a month. Seven hundred and twenty dollars a year. Run four and it is nine hundred and sixty.
The upper tiers are where it stops being a rounding error. Claude Max starts at one hundred dollars a month. ChatGPT Pro runs one hundred at the lower tier and two hundred at the upper. Google AI Ultra sits at two hundred, itself reduced from two hundred and fifty earlier this year. Keep two of those running because the release everyone is discussing is only good on the expensive plan, and you have committed somewhere between twenty-four hundred and forty-eight hundred dollars for the year.
For a small business that is not a software expense. It is a contractor for a week. It is a year of hosting for every client site you manage. It competes with things that generate revenue rather than opinions.
The annual discount is a bet that you will not switch
There is a detail in the pricing worth reading closely, because it tells you how these companies expect you to behave.
Claude Pro is twenty dollars a month billed monthly. It is seventeen dollars a month if you pay for the year up front, which is two hundred dollars in a single charge. The discount comes to thirty-six dollars across twelve months. Every one of these companies sells some version of this trade.
Take the discount and switch in month three, which is precisely what an eleven-day release cycle encourages, and the arithmetic turns over. You paid two hundred dollars. You used about fifty of it. You threw away one hundred and fifty, and you are now paying twenty a month for the replacement on top of what you already spent.
Saving thirty-six dollars requires being right about which tool you want for a full year. In a market shipping something new every eleven days, that is a genuine wager, and the discount exists because the company has a confident view of how it resolves.
Picking a good enough tool early and staying put will beat, on cost, almost any strategy built around finding the best one.
The ceiling in your own house
Benchmarks measure general capability. They rank models against each other across a wide spread of tasks, and they do that honestly enough. What they cannot tell you is anything about your Tuesday.
A model that scores highest overall is the tallest ladder in the shop. If your ceiling is eight feet, the tallest ladder is not the one you want. You are paying for reach you will never use and hauling weight you do not need every time you take it out of the van.
The work most of us actually do is narrow. Drafting client emails that sound like you wrote them. Cleaning up a spreadsheet somebody else built badly. Fixing a WordPress template that broke on a Sunday. Writing captions that do not read as machine output. None of that appears on a leaderboard, and a model that gained four points on graduate-level mathematics did not get better at any of it.
The honest test is one task you do every week, run against the tool you already pay for. If a new model does that task better, you have learned something worth money. If it does not, you have also learned something, and it cost one afternoon rather than a year of them.
The prices move in both directions
There is a reason to look up occasionally, and it is financial rather than technical.
Anthropic cut the price of cache reads by seventy-five percent this month, from one dollar per million tokens down to twenty-five cents on its newest model. Caching covers the cost of re-reading context the tool has already seen, which for anyone running long sessions is a large share of the bill. That is a real reduction in a real invoice, and nobody was required to tell you about it.
More telling: Claude Sonnet 5 had a price increase already scheduled. Two dollars per million input tokens was announced as introductory, due to rise to three dollars on September 1. The pricing documentation now states plainly that the increase will not happen and the lower price is permanent.
Nobody sends an email announcing a price increase that was cancelled. There is no launch event for a bill that stayed the same. That sat in a footnote on a documentation page, and the only way to have it is to go and read the table yourself.
So the case for paying attention is not that you are missing a smarter model. It is that you may be paying an old price for something that quietly got cheaper. That argues for checking deliberately every few months. It does not argue for switching every eleven days.
Exploring is not the same as following
Following means the release calendar sets your schedule. Something ships, you read the announcement, you feel behind, you enter a card number. The labs decide when you spend your afternoon and your money.
Exploring means you set the schedule. Pick a date, once a quarter. Bring the same three tasks you actually do. Run them against whatever is current, including the tool you already use, and give the incumbent a fair hearing rather than assuming the new arrival wins. Then decide, and go back to work until the next date.
Keep one daily driver you know well. Knowing a tool deeply is worth more than a marginal gain on paper, because you have learned where it fails and built around that. A new model resets all of it, and you are back to not knowing which of its answers to trust.
Two tools is reasonable, one for writing and one for code, if those are genuinely different jobs for you. Five is a hobby. Thirty-three is a full time position that nobody is paying you for.
Where this leaves you
The release pace is not going to slow. The petition will not change it and neither will anybody’s fatigue. Eleven days is the environment now, and it is the environment you have to run a business inside.
What that requires of you is smaller than it looks. Know what you need the tool to do. Check the prices twice a year, in both directions. Test on your own work rather than somebody’s benchmark. Change when the evidence says to change, and not because a company had a launch.
A model that has been working for you for six months has something no announcement can offer, which is six months of evidence.
Forward → Upward ↑ Onward ↗︎
Mstimaj
Sources and Further Reading
- Model fatigue sets in as AI labs race to roll out new versions at frenetic pace, CNBC.
- AI labs face model fatigue as breakneck release cycles take their toll, Crypto Briefing.
- Plans and Pricing, Claude by Anthropic.
- Pricing, Claude Platform Docs.
- Introducing Claude Fable 5.1 and Claude Mythos 5.1, Anthropic.
- Local AI Model Fatigue: Why One Setup Beats Chasing Every Release, MindStudio.
Join the Conversation
Share your thoughts and connect with other readers