Back to Articles

vLLM vLLM released 0.17.1 and fixed key patches for the inference backend, allowing patches such as TRTLLM MoE, Mamba/Qwen3.5 cache, and MTP processing to be implemented centrally

Found 1 related articles

Recommended Tools

More