The latest vLLM release, version 0.30.1rc0, brings significant enhancements for AMD GPU users. The update adds kernel mirrors for the MI355 accelerator, supporting both dense and Mixture of Experts (MoE) model architectures. These kernels are designed to maximize performance on AMD’s latest hardware, improving inference efficiency for large-scale machine learning workloads. The release notes, published on the vLLM GitHub repository, highlight contributions from AMD engineers and OpenAI’s Codex tool.
For developers and researchers working with AMD GPUs, this update removes a critical bottleneck. Prior to this release, optimizing model inference on MI355 hardware required manual configuration or third-party tools. The new kernels automate this process, reducing setup complexity and unlocking hardware-specific optimizations. This is particularly impactful for organizations relying on AMD’s ROCm ecosystem, which has historically lagged behind NVIDIA’s CUDA in AI infrastructure support. The addition of MoRI (Mixture of Experts Runtime Inference) kernels further extends vLLM’s compatibility, enabling efficient execution of sparse models that activate only subsets of parameters during inference.
The ability to run dense models (e.g., standard transformers) and MoE architectures like Google’s Switch Transformer on MI355 GPUs opens new possibilities for cost-effective scaling. MoE models, known for their parameter efficiency, are now more accessible to teams without dedicated NVIDIA infrastructure. This could accelerate experimentation with large language models (LLMs) in environments where AMD hardware is preferred for cost or supply chain reasons.
The release is available through Mina Labs, where developers can deploy these optimized kernels at a competitive rate. The cost for generating a single image using the updated vLLM stack is 8. Mina Labs’ platform abstracts much of the hardware complexity, allowing users to focus on model development rather than infrastructure tuning. The service integrates seamlessly with existing ML workflows, supporting both Python-based APIs and containerized deployments.
Practical applications for this release span a range of domains. Researchers training or deploying MoE models for natural language processing, computer vision, or multimodal tasks will benefit from reduced latency and memory usage. Enterprises running inference workloads on AMD GPUs can achieve faster throughput, improving response times for applications like real-time translation or recommendation systems. Additionally, the ROCm support aligns with broader industry trends toward hardware diversity, reducing vendor lock-in concerns for AI teams.
The MI355’s NVFP4 (NVIDIA Floating Point 4-bit) precision support also merits attention. While not a new feature per se, the kernel mirrors ensure that vLLM users can leverage low-precision arithmetic without sacrificing model accuracy. This is critical for edge deployments or scenarios where computational resources are constrained. Developers can now balance speed, memory, and precision more effectively, tailoring models to specific hardware profiles.
Looking ahead, this release signals a maturation of AMD’s AI ecosystem. As more frameworks adopt ROCm, the gap between NVIDIA and AMD in AI infrastructure narrows. For teams evaluating hardware options, the updated vLLM release provides a tangible reason to consider AMD accelerators, especially for workloads prioritizing cost efficiency and open-source compatibility. The integration with Mina Labs simplifies adoption, offering a turnkey solution for those seeking to harness MI355’s capabilities without in-house ROCm expertise.
In summary, vLLM 0.30.1rc0 is a targeted update with broad implications for AMD GPU users. By optimizing kernel performance and expanding MoE support, it enhances the viability of AMD hardware in production AI systems. Developers can now push the boundaries of model size and complexity while maintaining control over infrastructure costs.
MINA LABS
Start creating free