DeepSeek V4.1-Flash replaces V4 Pro routing with cheaper multimodal API access

DeepSeek V4.1-Flash release image from the official API documentation.DeepSeek API Docs
DeepSeek V4.1-Flash release image from the official API documentation.DeepSeek API Docs
AI & Automation

DeepSeek V4.1-Flash is live with native image understanding, lower token rates, and a September 14 routing change that moves V4 Pro requests to the new model.

DeepSeek has put V4.1-Flash into production on its API and is using it to phase out V4 Pro routing. The new model adds native multimodal input, a new asymmetric Causal Encoder–Decoder architecture, and materially lower prices for long-running agent workloads.

DeepSeek V4.1-Flash is live on the API

The official release identifies V4.1-Flash as an 8B-active-input, 16B-active-output model built on a 552B-parameter backbone. DeepSeek says native visual understanding is part of the architecture, rather than an add-on to the earlier V4 Flash Vision experiment. The model is aimed at coding, terminal, computer-use, and long-horizon tasks.

DeepSeek’s release says V4.1-Flash went live on September 10, 2026. The legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names remain temporarily compatible, but route to V4.1-Flash. The official page also links the released model artifacts and technical report.

September 14 changes the V4 Pro path

Starting at 04:00 UTC on September 14, all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates. DeepSeek describes this as a phase-out that will continue until V4.1-Pro launches. That is a routing and billing change for existing integrations, even when their code continues to use the V4 Pro model name.

The timing matters for teams with scheduled jobs or cost dashboards. A workload that depends on V4 Pro’s behavior should pin and re-evaluate its model assumptions instead of treating the compatibility route as a permanent equivalence. DeepSeek’s separate API documentation says V4 Pro service continues after September 14, so the safest reading is that the endpoint remains available while requests are served under the new routing policy.

Concept illustration: September 14 changes the V4 Pro path
AI-generated illustration

Lower prices target agent cache costs

DeepSeek attributes the efficiency gain partly to compressed KV caching. The official release says V4.1-Flash needs roughly one quarter of the previous Flash generation’s KV cache, which directly affects repeated-context and agentic workloads.

OpenRouter lists V4.1-Flash at $0.15 per million input tokens and $0.60 per million output tokens, with a 1,048,576-token context window and up to 1,048,576 output tokens. Its model page independently reports live access through multiple providers. Exact billing on DeepSeek’s own API still depends on its peak/off-peak schedule; the release says off-peak rates remain 50% of peak rates.

What to check in an existing integration

Before switching production traffic, compare tool calls, vision inputs, structured outputs, latency, and long-context completion behavior against the workload’s current V4 Pro baseline. Also inspect cost attribution after the September 14 UTC transition: compatibility aliases and explicit V4.1-Flash requests may now represent the same underlying service at different historical price assumptions.

The immediate change is clear: DeepSeek’s cheaper multimodal Flash tier is now the default destination for V4 Pro routing, while V4.1-Pro remains the next stated milestone.

Sources

From reading to doing

Try the related loot

Put six hosted Workers AI models behind Cloudflare AI Search

Open loot