184 days at OpenAI, 61 at Anthropic, 92 for a Codex variant. Retirement notice is a contract clause, and your teams are the ones absorbing it.
GPT-5.6 and Kimi K3 make 1M-token context the norm. Cost, recall, and hybrid architecture: what should actually change in your RAG.
A long-running agent can multiply your inference bill. Four concrete levers to regain control of token budgets in production.
Enterprise LLM spend doubled in 6 months. Semantic caching, model routing, and quantization can cut costs by 60% without touching output quality.