Abstrakte KI-Visualisierung fuer ein neues Sprachmodell

DeepSeek V4.1 Flash: New Encoder-Decoder Architecture and Open Weights

On September 10, 2026, DeepSeek unveiled V4.1 Flash, a multimodal language model that introduces a new architecture and sharply cuts API prices.

A new causal encoder-decoder architecture

The company calls V4.1 Flash the smallest model in a new family. It uses a “Causal Encoder–Decoder” design with a 552-billion-parameter mixture-of-experts core, of which DeepSeek says only 8 billion parameters are active on input and 16 billion on output. Compared with the previous generation, the model reportedly needs just a quarter of the HBM and an eighth of the SSD storage for its KV cache.

Multimodal, open weights, low prices

V4.1 Flash understands images natively and is reachable through the deepseek-flash endpoint with a one-million-token context window. DeepSeek publishes the weights as an open-weight model on Hugging Face, along with a technical report. Prices dropped at launch: during peak hours the price list quotes input at $0.30 per million tokens (cache miss) and output at $1.20, with off-peak rates set at half of that.

Performance claims come with caveats

DeepSeek points to “tests by multiple parties” that place V4.1 Flash ahead of the outgoing V4-Pro on performance, cost, speed and total runtime – vendor figures that have not yet been independently verified. From September 14, 2026, DeepSeek plans to route V4-Pro requests automatically to the new model until a V4.1-Pro variant arrives.

See also  Orion Browser for Linux: New Beta Update

Sources: DeepSeek – Introducing DeepSeek-V4.1-Flash and the announcement in the DeepSeek API Docs.

Leave a Comment

Your email address will not be published. Required fields are marked *

Mastodon
Scroll to Top