CVE-2025-49847
llama.cpp is an inference of several LLM models in C/C++. Prior to version b5662, an attacker‐supplied GGUF model vocabulary can trigger a buffer overflow in llama.cpp’s vocabulary‐loading code. Specifically, the helper _try_copy in llama.cpp/src/vocab.cpp: llama_vocab::impl::token_to_piece() casts a very large size_t token length into an int32_t, causing the length check (if (length < (int32_t)size)) to be bypassed. As a result, memcpy is still called with that oversized size, letting a malicious model overwrite memory beyond the intended buffer. This can lead to arbitrary memory corruption and potential code execution. This issue has been patched in version b5662.
CVSS
- Versión: 3.1
- Vector: CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H
- Puntuación base: 8.8
Probabilidad de explotación (EPSS)
- Probabilidad de explotación en los próximos 30 días: 0.52%
- Percentil entre todas las CVEs puntuadas: 42
- Fecha de la puntuación: 6/10/2026
EPSS (Exploit Prediction Scoring System, de FIRST) estima la probabilidad de que una vulnerabilidad sea explotada en 30 días. Complementa a CVSS (impacto) y a CISA KEV (explotación confirmada).
🎯 Técnicas ATT&CK
Cómo se explota esta vulnerabilidad y qué consigue el atacante, en el lenguaje de MITRE ATT&CK.
- Explotación
T1203Exploitation for Client Executionexecution85 % - Impacto principal
T1059Command and Scripting Interpreterexecution80 % - Impacto secundario
T1565Data Manipulationimpact70 %
UI:R requiere carga de modelo GGUF malicioso por usuario. Buffer overflow en vocab-loading causa corrupción de memoria y ejecución potencial de código arbitrario.
Inferido por nuestro agente de análisis a partir de la descripción oficial, el vector CVSS y la CWE, y comprobado por un supervisor. Puede contener errores.
🛡️ Mitigaciones ATT&CK que cubren estas técnicas
Tecnologías afectadas (1)
CWE
- CWE-119, CWE-195
Referencias
JSON original (NVD)
Mostrar
{
"id": "CVE-2025-49847",
"cveTags": [],
"metrics": {
"ssvcV203": [
{
"source": "134c704f-9b21-4f2e-91b3-4a467353bcc0",
"ssvcData": {
"id": "CVE-2025-49847",
"role": "CISA Coordinator",
"options": [
{
"exploitation": "poc"
},
{
"automatable": "no"
},
{
"technicalImpact": "total"
}
],
"version": "2.0.3",
"timestamp": "2025-06-18T13:40:43.172535Z"
}
}
],
"cvssMetricV31": [
{
"type": "Secondary",
"source": "security-advisories@github.com",
"cvssData": {
"scope": "UNCHANGED",
"version": "3.1",
"baseScore": 8.8,
"attackVector": "NETWORK",
"baseSeverity": "HIGH",
"vectorString": "CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"integrityImpact": "HIGH",
"userInteraction": "REQUIRED",
"attackComplexity": "LOW",
"availabilityImpact": "HIGH",
"privilegesRequired": "NONE",
"confidentialityImpact": "HIGH"
},
"impactScore": 5.9,
"exploitabilityScore": 2.8
}
]
},
"affected": [
{
"source": "security-advisories@github.com",
"affectedData": [
{
"vendor": "ggml-org",
"product": "llama.cpp",
"versions": [
{
"status": "affected",
"version": "< b5662"
}
]
}
]
}
],
"published": "2025-06-17T20:15:32.437",
"references": [
{
"url": "https://github.com/ggml-org/llama.cpp/commit/3cfbbdb44e08fd19429fed6cc85b982a91f0efd5",
"tags": [
"Patch"
],
"source": "security-advisories@github.com"
},
{
"url": "https://github.com/ggml-org/llama.cpp/security/advisories/GHSA-8wwf-w4qm-gpqr",
"tags": [
"Mitigation",
"Vendor Advisory"
],
"source": "security-advisories@github.com"
}
],
"vulnStatus": "Analyzed",
"weaknesses": [
{
"type": "Secondary",
"source": "security-advisories@github.com",
"description": [
{
"lang": "en",
"value": "CWE-119"
},
{
"lang": "en",
"value": "CWE-195"
}
]
}
],
"descriptions": [
{
"lang": "en",
"value": "llama.cpp is an inference of several LLM models in C/C++. Prior to version b5662, an attacker‐supplied GGUF model vocabulary can trigger a buffer overflow in llama.cpp’s vocabulary‐loading code. Specifically, the helper _try_copy in llama.cpp/src/vocab.cpp: llama_vocab::impl::token_to_piece() casts a very large size_t token length into an int32_t, causing the length check (if (length < (int32_t)size)) to be bypassed. As a result, memcpy is still called with that oversized size, letting a malicious model overwrite memory beyond the intended buffer. This can lead to arbitrary memory corruption and potential code execution. This issue has been patched in version b5662."
},
{
"lang": "es",
"value": "llama.cpp es una inferencia de varios modelos LLM en C/C++. Antes de la versión b5662, un vocabulario de modelo GGUF proporcionado por un atacante podía provocar un desbordamiento de búfer en el código de carga de vocabulario de llama.cpp. Específicamente, el asistente _try_copy en llama.cpp/src/vocab.cpp: llama_vocab::impl::token_to_piece() convierte una longitud de token size_t muy grande en un int32_t, lo que provoca que se omita la comprobación de longitud (si (length < (int32_t)size)). Como resultado, se sigue llamando a memcpy con ese tamaño excesivo, lo que permite que un modelo malicioso sobrescriba la memoria más allá del búfer previsto. Esto puede provocar corrupción de memoria arbitraria y la posible ejecución de código. Este problema se ha corregido en la versión b5662."
}
],
"lastModified": "2026-06-17T09:32:00.730",
"configurations": [
{
"nodes": [
{
"negate": false,
"cpeMatch": [
{
"criteria": "cpe:2.3:a:ggml:llama.cpp:*:*:*:*:*:*:*:*",
"vulnerable": true,
"matchCriteriaId": "CD259F6A-4B43-4B07-83A5-544F900CD023",
"versionEndExcluding": "b5662"
}
],
"operator": "OR"
}
]
}
],
"sourceIdentifier": "security-advisories@github.com"
}