Claude Code Falls Back to the Previous Model When the API Refuses Your Default
Claude Code 2.1.286 stops turns from failing outright when the Anthropic API refuses the model a default or alias resolves to: it now retries once on the previous model of the same tier. Fallback retries run at standard speed when the fallback cannot run fast, and notices now say when a fallback shrinks the context window from 1M to 200K tokens.
Key Takeaways
- When the API refuses the resolved model, Claude Code now retries once on the previous model of the same tier instead of failing every turn.
- Fallback retries run at standard speed when the fallback model cannot run fast, with a one-time notice in interactive sessions.
- Notices now warn when a fallback drops the context window from 1M to 200K tokens.
- A single retry limit caps a failing call at 14 requests with default settings.
- On Bedrock and Vertex AI, 2.1.285 switches to an older available model of the same tier if an admin removes the default.
- Earlier retry storms of up to 21 requests on failing streams were removed by sharing one budget with the non-streaming fallback.
One refusal no longer breaks every turn
Previously, if the Anthropic API refused the model that a default setting or a model alias resolved to, every turn in the session could fail. Claude Code 2.1.286 changes that: it now retries once on the previous model of the same tier. A developer who selected the newest model through an alias keeps working on the prior release instead of being blocked. The related 2.1.285 release applied similar logic to Amazon Bedrock and Vertex AI, switching to an older available model of the same tier when an admin removes access to the default, with session titles and summaries falling back along with it.
Fast mode and fallback retries
Refusal and --fallback-model retries also failed when the fallback model could not run in fast mode. They now run at standard speed instead, and interactive sessions show a one-time notice explaining the change.
Clear notices about context loss
A fallback can move a session from a 1M-token context window to 200K. The model fallback notice and the autocompact-thrashing error now say so when that happens, which explains why a long session may suddenly compact more often. Developers who do not want a smaller window can pick a model with the larger limit.
Related reliability changes
Retry behavior was also tightened: a single limit now covers a whole model call, so with default retry settings a failing call sends at most 14 requests. The 2.1.285 release similarly stopped a failing streamed request from being retried up to 21 times by sharing the retry budget with the non-streaming fallback.