vllm
8 records · 8 with a public proof-of-concept
Disclosed vulnerabilities where the NVD names vllm as an affected vendor, highest CVSS first.
- MEDIUM 6.5CVE-2026-105757public PoC
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compila…
vllm
- MEDIUM 6.5CVE-2026-105756public PoC
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, OpenAI-compatible request models accept a non-empty cache_salt value without enforcing the character and length restrictions required by the IPCCacheServerKey consumer in LMCac…
vllm
- MEDIUM 6.5CVE-2026-105754public PoC
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the features.mm_…
vllm
- MEDIUM 6.5CVE-2026-105753public PoC
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engine receive…
vllm
- MEDIUM 5.3CVE-2026-105758public PoC
vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the media_io_kwargs.video.max_frames and media_io_kwargs.video.fps fields witho…
vllm
- MEDIUM 5.3CVE-2026-103241public PoC
A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of the component Gemma4UnifiedParser. Executing a manipulation can lead to denial of service. The attack may be l…
vllm
- MEDIUM 4.2CVE-2026-105755public PoC
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's query_key value from the caller-controlled X-Request-Id header. A concurrent request th…
vllm
- LOW 3.1CVE-2026-105752public PoC
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the continuation …
vllm