Vulnerabilidades
Resumen — últimos 7 días
Vulnerabilidades nuevas2855▼ 333 respecto a la semana anterior
Críticas / altas1381▼ 36 respecto a la semana anterior
Nueva explotación activa (KEV)4▼ 5 respecto a la semana anterior
Sin puntuar (sin CVSS)296▼ 213 respecto a la semana anterior
109 resultados, ordenados por fecha de publicación (más recientes primero)
| CVE | Estado | Severidad | EPSS | Explotación activa | Tecnologías afectadas | Publicada ▼ | Modificada | Descripción |
|---|---|---|---|---|---|---|---|---|
| Analizada | Alta (8.7) | 0.63% | — | Vllm | 21/9/2026 | 29/9/2026 | vLLM versions through 0.29.0 contain a denial of service vulnerability in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests with incomplete kv_transfer_params dictionary entries to trigger an uncaught KeyError in EngineCore scheduling, causing the decode… | |
| Analizada | Baja (2.3) | 0.20% | — | Vllm | 19/9/2026 | 28/9/2026 | vLLM through 0.29.0 fails to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). Attackers can supply out-of-bounds token indices that corrupt logits memory of concurrent requests, causing different in-flight HTTP requests to return incorrect… | |
| Analizada | Media (6.3) | 0.24% | — | Vllm | 18/9/2026 | 28/9/2026 | vLLM through 0.29.0 contains a memory corruption vulnerability in the Triton _bincount_kernel where prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. Attackers can submit multimodal audio requests with tokens equal to vocabulary size, causing out-of-bounds… | |
| Analizada | Media (6.3) | 0.25% | — | Vllm | 18/9/2026 | 28/9/2026 | vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass validation, causing LogitBiasState to corrupt GPU logits state and allow concurrent requests… | |
| Analizada | Alta (8.7) | 0.54% | — | Vllm | 18/9/2026 | 28/9/2026 | vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context,… | |
| Analizada | Alta (8.7) | 0.76% | — | Vllm | 17/9/2026 | 28/9/2026 | vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts. | |
| En análisis | Media (6.5) | 0.55% | — | VllmAINvidia PynvvideocodecAI | 16/9/2026 | 30/9/2026 | vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice to VideoMediaIO even when startup configuration selected a software… | |
| En análisis | Media (6.5) | 0.69% | — | VllmAI | 16/9/2026 | 24/9/2026 | vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATION_S to the shared audio decoder. An unauthenticated client can therefore submit… | |
| Aplazada | Media (5.3) | 0.52% | — | Vllm-project VllmAI | 16/9/2026 | 22/9/2026 | A vulnerability was found in vllm-project vllm up to 0.29.0. Affected by this issue is some unknown functionality of the file vllm/v1/sample/thinking_budget_state.py. The manipulation results in inefficient algorithmic complexity. It is possible to launch the attack remotely. The pull request to fix this issue awaits… | |
| Aplazada | Media (6.9) | 0.70% | — | Vllm-project VllmAI | 16/9/2026 | 16/9/2026 | A vulnerability was found in vllm-project vLLM 0.26.0/0.27.0. Affected is the function MoRIIOConnectorScheduler.request_finished/MoRIIOConnectorWorker.get_finished/MoRIIOWrapper._handle_release_message of the file vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py of the component MoRIIO… | |
| Aplazada | Baja (2.1) | 0.53% | — | Vllm-project VllmAI | 15/9/2026 | 15/9/2026 | A vulnerability was determined in vllm-project vLLM up to 0.27.1. This affects an unknown part of the file /v1/chat/completions of the component Jinja Template Rendering. This manipulation of the argument chat_template causes resource consumption. The attack can be initiated remotely. The exploit has been publicly… | |
| Aplazada | Baja (1.9) | 0.16% | — | Vllm-project VllmAI | 14/9/2026 | 15/9/2026 | A security flaw has been discovered in vllm-project vLLM up to 0.29.0. The affected element is the function TiktokenTokenizer::new of the file rust/src/text/src/backend/hf/mod.rs of the component tiktoken vocab File Handler. The manipulation results in denial of service. The attack is only possible with local access.… | |
| Analizada | Alta (7.1) | 0.52% | — | Vllm | 12/9/2026 | 16/9/2026 | vLLM versions before 0.28.0 fail to validate audio sample rate headers in the transcription endpoint, allowing authenticated clients to bypass duration checks. Attackers can submit forged FLAC headers with inflated sample rates to trigger excessive memory allocation and crash the API server process affecting all… | |
| Analizada | Alta (8.5) | 0.31% | — | Vllm | 12/9/2026 | 16/9/2026 | vLLM before 0.28.0 contains a remote code execution vulnerability in the LlavaOnevision2 processor loader that ignores the trust_remote_code parameter when loading remote processor classes. Attackers can craft a malicious model with arbitrary code in processing_llava_onevision2.py that executes with vLLM process… | |
| Analizada | Media (6.9) | 0.20% | — | Vllm | 12/9/2026 | 24/9/2026 | vLLM versions >=0.10.2 and <0.28.0 do not apply any audio decode-size or duration limit when extracting audio from video input for NanoNemotronVL models. In nano_nemotron_vl.py, _extract_audio_from_videos calls load_audio_pyav(BytesIO(video_bytes)) without the max_duration_s or max_decode_bytes parameters, so neither… | |
| Analizada | Alta (7.5) | 0.73% | — | Vllm | 28/8/2026 | 29/9/2026 | vLLM up to and including 0.17.0 allows remote attackers to cause a Denial of Service via memory exhaustion. The AsyncMediaIO.fetch_audio and AsyncMediaIO.fetch_image functions in multimodal/inputs.py fetch user-supplied media URLs using aiohttp and call r.read() without enforcing a maximum response size, allowing an… | |
| Analizada | Media (6.9) | 0.61% | — | Vllm | 25/8/2026 | 29/9/2026 | vLLM before 0.27.0 fails to properly classify DeepStream as a GPU backend and omits pixel-limit enforcement in its decode path. Unauthenticated attackers can activate DeepStream at request time to initialize the process-wide GPU decode pool and submit video that bypasses resource controls, causing partial denial of… | |
| Analizada | Media (6.5) | 0.41% | — | Vllm | 17/8/2026 | 28/9/2026 | vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the MiMoV2OmniMultiModalProcessor in vllm/transformers_utils/processors/mimo_v2_omni.py passes attacker-controlled image and audio strings through _fetch_image, requests.get, and Image.open instead of MediaConnector, bypassing… | |
| Analizada | Media (4.3) | 0.37% | 💥 PoC | Vllm | 17/8/2026 | 2/10/2026 | vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices, token_ids, prompt_logprobs, logprobs.content, top_logprobs, and routed_experts… | |
| En análisis | Media (6.3) | 0.40% | — | VllmAI | 13/8/2026 | 9/9/2026 | vLLM is an inference and serving engine for large language models. From 0.20.2rc0 until 0.26.0, safe_load_prompt_embeds in vllm/renderers/embed_utils.py uses torch.sparse.check_sparse_tensor_invariants, whose process-global save, enable, and restore state can be raced by concurrent prompt_embeds parts submitted to… | |
| Analizada | Media (6.5) | 0.58% | — | Vllm | 13/8/2026 | 28/9/2026 | vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and… | |
| Analizada | Media (5.3) | 0.41% | — | Vllm | 13/8/2026 | 28/9/2026 | vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x * 2 * d in activation_kernels.cu can cause act_and_mul_kernel to consume another batched user's input, allowing a request processed in the same inference batch to receive a partial or complete copy of… | |
| Analizada | Media (5.3) | 0.52% | — | Vllm | 13/8/2026 | 28/9/2026 | vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the structured_outputs.regex parameter in vllm/v1/structured_output/backend_lm_format_enforcer.py is passed to lmformatenforcer.RegexParser without compile_regex_with_timeout or validation in… | |
| Analizada | Media (5.3) | 0.41% | — | Vllm | 13/8/2026 | 2/10/2026 | vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the validation_exception_handler in vllm/entrypoints/openai/server_utils.py converts FastAPI RequestValidationError objects with str(exc), and sanitize_message in vllm/entrypoints/utils.py does not remove traceback-style file paths,… | |
| Analizada | Media (6.8) | 0.10% | — | Intel Vllm Hardware | 11/8/2026 | 24/9/2026 | Improper input validation for some vLLM Hardware Plugin for Intel(R) Gaudi(R) software before version 0.16.0 within Ring 3: User Applications may allow a denial of service. Authorized adversary with an authenticated user combined with a low complexity attack may enable denial of service. This result may potentially… |