Vllm
Vllm: vulnerabilidades y CVE
Vllm tiene 87 vulnerabilidades publicadas, 68 de ellas en los últimos 12 meses. 8 son críticas y 0 figuran en el catálogo de explotación activa de CISA.
CVE87
Últimos 12 meses68
Críticas8
Explotadas activamente0
Todas las vulnerabilidades en el catálogo →⭐ Seguir esta tecnología
Últimas vulnerabilidades
| CVE | Severidad | EPSS | Explotación activa | Publicada | Descripción |
|---|---|---|---|---|---|
| CVE-2026-100654 | Alta (7.1) | 0.31% | — | 26 sept 2026 | vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints but validates only that the values are integers, not that each token id is… |
| CVE-2026-100653 | Alta (8.3) | 0.28% | — | 26 sept 2026 | vLLM is an inference and serving engine for large language models. In versions from 0.22.1 through 0.28.0, the operator-supplied model revision pin (--revision / --code-revision) is not propagated to several Hugging… |
| CVE-2026-100652 | Alta (8.2) | 0.33% | — | 26 sept 2026 | vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit… |
| CVE-2026-100651 | Alta (7.1) | 0.31% | — | 26 sept 2026 | vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload,… |
| CVE-2026-100650 | Alta (7.1) | 0.61% | — | 26 sept 2026 | vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the… |
| CVE-2026-100649 | Media (6.3) | 0.31% | — | 26 sept 2026 | vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select… |
| CVE-2026-100648 | Media (6.9) | 0.33% | — | 26 sept 2026 | vllm before 0.29.0 fails to enforce VLLM_MAX_AUDIO_CLIP_FILESIZE_MB limit in multimodal chat audio decoding, allowing unauthenticated clients to bypass file size restrictions. Attackers can submit oversized audio files… |
| CVE-2026-100647 | Media (6.9) | 0.31% | — | 26 sept 2026 | vLLM versions before 0.29.0 contain a denial-of-service vulnerability in the cache_salt parameter accepted on OpenAI-compatible and Anthropic API endpoints, which lacks maximum length validation and is processed on the… |
| CVE-2026-94627 | Alta (8.7) | 0.63% | — | 21 sept 2026 | vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger… |
| CVE-2026-94626 | Alta (8.7) | 0.63% | — | 21 sept 2026 | vLLM through 0.29.0 fails to validate the tp_size parameter in kv_transfer_params on OpenAI-compatible completion endpoints, allowing attackers to allocate unbounded memory. Attackers can supply arbitrary tp_size values… |
| CVE-2026-94625 | Media (6.9) | 0.52% | — | 21 sept 2026 | vLLM through 0.29.0 contains a resource exhaustion vulnerability in MooncakeConnector where rejected prefill requests create ownerless transfer placeholders that are never reclaimed. Attackers can send rejected requests… |
| CVE-2026-94624 | Alta (8.7) | 0.63% | — | 21 sept 2026 | vLLM through 0.29.0 contains a denial of service vulnerability in P2P KV offloading when OffloadingConnector is configured with TieringOffloadingSpec and a peer-to-peer secondary tier. Attackers can supply arbitrary… |
| CVE-2026-94623 | Alta (8.7) | 0.63% | — | 21 sept 2026 | vLLM through 0.29.0 contains a denial of service vulnerability in the NIXL connector's prefix caching implementation that fails to properly validate block counts across multi-prompt completion requests in prefill/decode… |
| CVE-2026-94622 | Alta (8.7) | 0.63% | — | 21 sept 2026 | vLLM versions through 0.29.0 contain a denial of service vulnerability in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests with incomplete… |
| CVE-2026-93989 | Baja (2.3) | 0.20% | — | 19 sept 2026 | vLLM through 0.29.0 fails to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). Attackers can supply out-of-bounds token indices that corrupt… |
| CVE-2026-93841 | Media (6.3) | 0.24% | — | 18 sept 2026 | vLLM through 0.29.0 contains a memory corruption vulnerability in the Triton _bincount_kernel where prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. Attackers… |
| CVE-2026-93840 | Media (6.3) | 0.25% | — | 18 sept 2026 | vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary… |
| CVE-2026-93592 | Alta (8.7) | 0.54% | — | 18 sept 2026 | vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single… |
| CVE-2026-93436 | Alta (8.7) | 0.76% | — | 17 sept 2026 | vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust… |
| CVE-2026-69147 | Media (6.5) | 0.55% | — | 16 sept 2026 | vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and… |
| CVE-2026-57173 | Media (6.5) | 0.69% | — | 16 sept 2026 | vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing… |
| CVE-2026-90555 | Alta (7.1) | 0.52% | — | 12 sept 2026 | vLLM versions before 0.28.0 fail to validate audio sample rate headers in the transcription endpoint, allowing authenticated clients to bypass duration checks. Attackers can submit forged FLAC headers with inflated… |
| CVE-2026-90553 | Alta (8.5) | 0.31% | — | 12 sept 2026 | vLLM before 0.28.0 contains a remote code execution vulnerability in the LlavaOnevision2 processor loader that ignores the trust_remote_code parameter when loading remote processor classes. Attackers can craft a… |
| CVE-2026-90554 | Media (6.9) | 0.20% | — | 12 sept 2026 | vLLM versions >=0.10.2 and <0.28.0 do not apply any audio decode-size or duration limit when extracting audio from video input for NanoNemotronVL models. In nano_nemotron_vl.py, _extract_audio_from_videos calls… |
| CVE-2026-37237 | Alta (7.5) | 0.73% | — | 28 ago 2026 | vLLM up to and including 0.17.0 allows remote attackers to cause a Denial of Service via memory exhaustion. The AsyncMediaIO.fetch_audio and AsyncMediaIO.fetch_image functions in multimodal/inputs.py fetch user-supplied… |
| CVE-2026-78684 | Media (6.9) | 0.61% | — | 25 ago 2026 | vLLM before 0.27.0 fails to properly classify DeepStream as a GPU backend and omits pixel-limit enforcement in its decode path. Unauthenticated attackers can activate DeepStream at request time to initialize the… |
| CVE-2026-73560 | Media (6.5) | 0.41% | — | 17 ago 2026 | vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the MiMoV2OmniMultiModalProcessor in vllm/transformers_utils/processors/mimo_v2_omni.py passes attacker-controlled image and audio… |
| CVE-2026-71486 | Media (4.3) | 0.47% | — | 17 ago 2026 | vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose… |
| CVE-2026-73557 | Media (6.3) | 0.40% | — | 13 ago 2026 | vLLM is an inference and serving engine for large language models. From 0.20.2rc0 until 0.26.0, safe_load_prompt_embeds in vllm/renderers/embed_utils.py uses torch.sparse.check_sparse_tensor_invariants, whose… |
| CVE-2026-73559 | Media (6.5) | 0.58% | — | 13 ago 2026 | vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded… |
🎯 Cómo se explota (técnicas ATT&CK)
Número de CVE de esta tecnología asignadas a cada técnica de explotación o de impacto principal.