The Chinese company DeepSeek has unveiled an updated version of its DeepSeek-V4-Flash-0731 model, demonstrating that further improvements in artificial intelligence performance do not necessarily require increasingly larger models. The new version features a significantly enhanced architecture and size thanks to new post-training techniques, and it even outperforms the larger DeepSeek-V4-Pro model in independent tests.

DeepSeek-V4-Flash-0731 builds upon the April preliminary version of the V4-Flash model. According to the company's documentation, the architecture and size of the model have remained unchanged, with a focus on new fine-tuning after initial training. This step has resulted in a significant increase in capabilities, highlighting the crucial role that post-training methods play alongside the sheer number of parameters.

The model utilizes a Mixture-of-Experts (MoE) architecture and activates only a portion of its parameters for each token. The official checkpoint released by DeepSeek has approximately 304 billion parameters, including a module for speculative decoding. The model is also available as an open-weight solution through the Hugging Face platform.

One of the most notable features is its exceptionally long context window. DeepSeek claims support for up to one million tokens on input and a maximum output length of 384,000 tokens. The model supports both reasoning mode and standard mode, as well as tool calling and other functions necessary, for example, when creating more autonomous AI agents.

Smaller Model Outperforms Larger Sibling

The results from the independent company Artificial Analysis are among the most interesting aspects of this update. In reasoning mode, DeepSeek-V4-Flash-0731 achieved a score of 50 points on the Artificial Analysis Intelligence Index. This places it among the most powerful current models and, according to the evaluator, significantly above the previous V4-Flash version.

The updated Flash model has also demonstrated its ability to outperform the larger DeepSeek-V4-Pro in tests, which is particularly significant from an efficiency perspective. The development of AI is increasingly shifting away from the question of "how to create the largest model" towards the question of "how to achieve the highest level of intelligence with available computing power."

DeepSeek is also putting pressure on competitors through pricing. Artificial Analysis reports a price of approximately $0.14 per million input tokens and $0.28 per million output tokens. In a combined scenario that includes cache usage, the cost is significantly lower than many competing models.

The combination of performance and price has placed DeepSeek-V4-Flash-0731 on what is known as the Pareto frontier by Artificial Analysis – meaning it is among the models for which it is not possible to find a simultaneously more intelligent and cheaper alternative based on the metrics being tracked.

Long Context and Faster Generation

An important aspect of the model is its optimization for working with long contexts. The V4 family architecture uses mechanisms that reduce the amount of computation and memory required when processing large inputs. This is crucial, for example, when working with large documents, databases, lengthy conversations, or in autonomous tasks where the model gradually accumulates a large amount of information.

The released version also includes speculative decoding technology. A smaller auxiliary module proposes several subsequent tokens in advance, and the main model can then verify them simultaneously. This results in faster generation without requiring significant changes to the underlying model itself. Hugging Face lists a checkpoint size of approximately 304 billion parameters for the released version.

Artificial Analysis measured over one hundred output tokens per second for the model, which ranks it among the fast reasoning models in this performance category.

The competition is shifting from size to efficiency

DeepSeek-V4-Flash-0731 demonstrates a broader trend that is increasingly prevalent in the generative AI industry. Developers are no longer competing solely on the absolute performance of the largest models, but also on how much intelligence they can offer for one dollar and with what hardware requirements.

This is particularly important for AI agents. These agents can consume millions of tokens when working with documents, programming, customer requests, or corporate databases during long-term task resolution. Even a relatively small difference in the price of one million tokens therefore translates into significant differences in operating costs when deployed on a large scale.

DeepSeek also has another advantage in the form of available model weights. This means that companies are not solely reliant on cloud APIs and can deploy the model on their own infrastructure, which can be interesting, for example, where it is important to keep sensitive data within the organization. DeepSeek released the model as part of its V4 collection on Hugging Face.

The V4-Flash update is therefore not only interesting because of another advancement in benchmark tables. It shows that better training and optimization can have a greater impact than simply increasing the number of parameters. And it is precisely the combination of high performance, speed, long context, open weights, and low price that may be more important for the further development of the AI market than the race to create the absolute largest model.

gnews.cz - tk