Chinese artificial intelligence company DeepSeek has launched DeepSeek-V4.1-Flash, a new multimodal model designed to improve inference speed, throughput and cost efficiency while supporting large-scale AI workloads.
Released on September 10, V4.1-Flash is the smallest model in DeepSeek's new architecture family and supports native visual understanding alongside text. The model is now available through the DeepSeek API and has also been released with open weights under an MIT licence.
V4.1-Flash is a 552-billion-parameter Mixture-of-Experts model built on a new Causal Encoder-Decoder architecture. Despite its overall size, DeepSeek says the architecture activates only 8 billion parameters while processing input and 16 billion while generating output, reducing the computational requirements associated with running the model.
A major focus of the release is reducing the infrastructure required for long-context and agentic AI workloads. According to DeepSeek, V4.1-Flash requires one-quarter of the high-bandwidth memory and one-eighth of the SSD storage for its key-value cache compared with the previous generation.
The model supports a context window of up to one million tokens and can process both text and images. These capabilities position it for applications involving coding, document analysis, visual understanding and AI agents that need to work across large amounts of information.
DeepSeek also said new pre-training methods and larger-scale reinforcement learning were used to improve the model's capabilities. The company claims its internal benchmark results place V4.1-Flash ahead of its V4-Pro model across several performance measures. These benchmark results are company-reported and can vary depending on testing conditions and workloads.
The release will also reshape DeepSeek's existing model lineup. V4-Flash and V4-Flash-Vision-Exp have been retired, with their API identifiers temporarily redirected to V4.1-Flash for compatibility. DeepSeek is also phasing out V4-Pro. From September 14, requests to V4-Pro are scheduled to be routed to V4.1-Flash until the company introduces V4.1-Pro.
DeepSeek has introduced revised API pricing alongside the launch, with off-peak usage priced at half the peak rate. The company said the more efficient architecture allows it to serve greater volumes of requests at lower cost.
The launch continues DeepSeek's push towards AI models designed to balance capability with lower inference and infrastructure requirements, as developers increasingly consider operating costs alongside model performance when deploying generative and agentic AI applications.
Disclaimer: This article may include information derived from interviews, press releases, public statements, research, company communications and other publicly available or third-party sources. Such material may be summarised, paraphrased or contextualised for journalistic and editorial purposes. All rights in third-party content remain with their respective owners.