DeepSeek’s new model turns memory savings into lower prices
Friday, September 11, 2026
DeepSeek has introduced V4.1-Flash, an AI model that the company says outperforms its previous V4-Pro while using less computing memory. Its attention-grabbing rate—less than a cent per million tokens, or small pieces of text processed by a model—applies when software reuses material the model has already seen; fresh input and generated answers cost more, while off-peak use is cheaper. That matters especially for AI agents: programs that work through long tasks and repeatedly return to earlier instructions, code, or documents. DeepSeek plans to phase out V4-Pro in favor of the new model, turning a more efficient design into added price pressure on rivals such as Anthropic and Z.ai.
Did you like the content?
