Vulnerabilidades

Resumen — últimos 7 días

Vulnerabilidades nuevas2856▼ 331 respecto a la semana anterior
Críticas / altas1383▼ 38 respecto a la semana anterior
Nueva explotación activa (KEV)4▼ 5 respecto a la semana anterior
Sin puntuar (sin CVSS)292▼ 217 respecto a la semana anterior
–

109 resultados, ordenados por fecha de publicación (más recientes primero)

CVEEstadoSeveridadEPSS Explotación activaTecnologías afectadasPublicada ▼Modificada Descripción
En análisisBaja (2.1)0.30%—Vllm-project VllmAI6/10/20266/10/2026
A security flaw has been discovered in vllm-project vLLM up to 0.31.0. This impacts the function get_token_bin_counts_and_mask of the file vllm/model_executor/layers/utils.py of the component Penalty Handler. Performing a manipulation results in denial of service. Remote exploitation of the attack is possible. The…
AplazadaBaja (2.1)0.30%—Vllm-project VllmAI6/10/20266/10/2026
A security vulnerability has been detected in vllm-project vLLM up to 0.31.0. This impacts the function conv_ssm_forward of the file vllm/model_executor/layers/mamba/mamba_mixer2.py of the component Completions Request Handler. The manipulation leads to out-of-bounds read. The attack is possible to be carried out…
En análisisMedia (5.3)0.30%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level media_io_kwargs field to select the GLMGA video backend and supply large values for the fps and max_frames options without a strict work ceiling. GLMGA constructs and deduplicates an attacker-sized…
En análisisMedia (5.9)0.33%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the Rust frontend's track_http_metrics middleware records the raw HTTP method token as a Prometheus label for requests reaching registered routes. An unauthenticated attacker can send unique arbitrary method tokens to unguarded routes…
En análisisMedia (5.3)0.31%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the media_io_kwargs.video.max_frames and media_io_kwargs.video.fps fields without enforcing server-side ceilings. An unauthenticated…
En análisisMedia (6.5)0.31%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compilation exception, padding produced by the ngram_gpu…
En análisisMedia (6.5)0.31%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, OpenAI-compatible request models accept a non-empty cache_salt value without enforcing the character and length restrictions required by the IPCCacheServerKey consumer in LMCache-MP. On deployments using the LMCache-MP connector, a…
En análisisMedia (4.2)0.20%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim's identifier can overwrite the cached…
En análisisMedia (6.5)0.27%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the features.mm_hashes field, ranges in the features.mm_placeholders field,…
En análisisMedia (6.5)0.43%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engine receiver cache never receives the payload if that request is…
En análisisBaja (3.1)0.21%—VllmAI5/10/20266/10/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the continuation prefix in the global unsalted cache namespace even when the…
AnalizadaMedia (5.5)0.71%—Vllm30/9/20266/10/2026
A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of the component Gemma4UnifiedParser. Executing a manipulation can lead to denial of service. The attack may be launched remotely. The exploit has been published and may be…
En análisisAlta (7.1)0.31%—VllmAI26/9/202630/9/2026
vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints but validates only that the values are integers, not that each token id is within the model vocabulary/logits range. When min_tokens > 0, the stop token ids are used as logits…
En análisisAlta (8.3)0.28%—VllmAI26/9/202630/9/2026
vLLM is an inference and serving engine for large language models. In versions from 0.22.1 through 0.28.0, the operator-supplied model revision pin (--revision / --code-revision) is not propagated to several Hugging Face artifact loads for the FunAudioChat and Tarsier2 architectures: the WhisperFeatureExtractor and…
AnalizadaAlta (8.2)0.36%—Vllm26/9/20266/10/2026
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids to trigger CUDA…
AnalizadaAlta (7.1)0.33%—Vllm26/9/20266/10/2026
vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload, vllm/entrypoints/serve/disagg/serving.py builds a multimodal EngineInput directly from the caller-supplied token_ids, and…
AnalizadaAlta (7.1)0.65%—Vllm26/9/20266/10/2026
vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the per-modality --limit-mm-per-prompt item limits). Across four ingress paths — the shared media-acquisition layer…
AnalizadaMedia (6.3)0.34%—Vllm26/9/20266/10/2026
vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video requests to exceed configured decoder limits and exhaust unaccounted…
AnalizadaMedia (6.9)0.35%—Vllm26/9/20266/10/2026
vllm before 0.29.0 fails to enforce VLLM_MAX_AUDIO_CLIP_FILESIZE_MB limit in multimodal chat audio decoding, allowing unauthenticated clients to bypass file size restrictions. Attackers can submit oversized audio files through chat endpoints to consume excessive memory and CPU resources during decoding.
AnalizadaMedia (6.9)0.35%—Vllm26/9/20266/10/2026
vLLM versions before 0.29.0 contain a denial-of-service vulnerability in the cache_salt parameter accepted on OpenAI-compatible and Anthropic API endpoints, which lacks maximum length validation and is processed on the single EngineCore scheduler thread. Unauthenticated attackers can send HTTP requests with…
AnalizadaAlta (8.7)0.63%—Vllm21/9/202629/9/2026
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts, causing orphaned KV…
AnalizadaAlta (8.7)0.63%—Vllm21/9/202629/9/2026
vLLM through 0.29.0 fails to validate the tp_size parameter in kv_transfer_params on OpenAI-compatible completion endpoints, allowing attackers to allocate unbounded memory. Attackers can supply arbitrary tp_size values in prefill/decode disaggregated deployments to exhaust memory and trigger kernel OOM-kill of the…
AnalizadaMedia (6.9)0.52%—Vllm21/9/202629/9/2026
vLLM through 0.29.0 contains a resource exhaustion vulnerability in MooncakeConnector where rejected prefill requests create ownerless transfer placeholders that are never reclaimed. Attackers can send rejected requests to exhaust sender task pools, causing valid requests to be delayed by up to 480 seconds while…
AnalizadaAlta (8.7)0.63%—Vllm21/9/202629/9/2026
vLLM through 0.29.0 contains a denial of service vulnerability in P2P KV offloading when OffloadingConnector is configured with TieringOffloadingSpec and a peer-to-peer secondary tier. Attackers can supply arbitrary remote host and port values in kv_transfer_params to create unreachable peer sessions that retain…
AnalizadaAlta (8.7)0.63%—Vllm21/9/202629/9/2026
vLLM through 0.29.0 contains a denial of service vulnerability in the NIXL connector's prefix caching implementation that fails to properly validate block counts across multi-prompt completion requests in prefill/decode disaggregated deployments. Attackers can trigger an assertion failure in…
Orbitaley — Vulnerabilidades