Mina LabsMINA LABS Start creating free
Blog / News
vLLM v0.31.0 adds DeepSeek-V4.1-Flash support

vLLM v0.31.0 adds DeepSeek-V4.1-Flash support

2026-10-05

vLLM released v0.31.0 on Monday, 5 October 2026. The release includes updates for DeepSeek-V4.1-Flash performance, including a new SM100 default for FlashMLA mega attention with the model’s NVFP4 compressed KV cache. Read the full announcement in the v0.31.0 release notes.

The release is substantial in scope. vLLM says it contains 717 commits from 307 contributors, including 96 contributors who are new to the project. The headline change is focused on serving DeepSeek-V4.1-Flash, with two performance-related pieces called out in the announcement: FlashMLA mega attention and DeepGEMM sparse MQA logits for the indexer.

For someone making things, this matters at the serving layer rather than the prompt or interface layer. vLLM is the infrastructure that helps applications run language models. Changes in attention handling, cache formats, and indexer computation can affect how a model behaves when it is deployed behind an application. That makes this release relevant to teams building assistants, research tools, automated workflows, and other products that need to call a model repeatedly.

The release specifically makes FlashMLA mega attention with the V4.1 NVFP4 compressed KV cache the default on SM100 hardware. The wording is important. This is not presented as a universal default for every machine. It is tied to SM100, so the practical value depends on the hardware used for serving. If your deployment matches that target, the release gives you an updated path to test rather than requiring you to assemble the configuration yourself.

The announcement also identifies DeepGEMM sparse MQA logits for the indexer. The source does not provide a benchmark, latency figure, or memory comparison in the supplied release summary. That means the useful next step is measurement on your own workload. Check response time, throughput, memory use, and output quality before changing a production deployment.

On Mina Labs, vLLM v0.31.0 is available to use. The listed price is 4 credits for one image. Because the release’s main changes concern model serving and hardware-specific execution, the important question is not only the listed price. It is whether the available setup matches the kind of generation or model workflow you are building and whether the SM100-specific defaults apply to your run.

We would use this release to test DeepSeek-V4.1-Flash in a production-shaped workflow. First, we would run a fixed set of prompts through the Mina Labs version and record latency, output consistency, and resource behavior. The test set would include short requests, long-context requests, and repeated calls that exercise the KV cache. That would show whether the new attention path is useful for the actual application instead of relying on the release description alone.

We would also compare the new setup with the configuration already used by an existing project. The comparison should keep prompts and generation settings constant. For a chat product, we would look at how the system handles concurrent sessions. For an internal research tool, we would focus on long inputs and repeatability. For an automated content pipeline, we would measure completed jobs and inspect the outputs for regressions.

The release is most useful today for builders who already have a reason to run DeepSeek-V4.1-Flash and access to compatible SM100 infrastructure. It offers a current vLLM base with the named performance paths enabled by default for that target. The source does not establish a universal speedup, so the right conclusion is practical: install or select v0.31.0, run your workload, and keep it if the measurements support the change.

Source: vLLM releases, https://github.com/vllm-project/vllm/releases/tag/v0.31.0 Make something with itMina Labs runs these models in your browser. Pay per generation, no subscription.