CVE-2026-72818
The URLS regular expression in nltk/tokenize/casual.py, compiled into TweetTokenizer.WORD_RE and applied by TweetTokenizer.tokenize, contains a naked-domain branch whose domain-label prefix [a-z0-9]+(?:[.\-][a-z0-9]+)* is unbounded. Input consisting of many alternating label separators can be partitioned in exponentially many ways, and because the branch also requires a trailing top-level domain that such input never supplies, the engine explores those partitions before failing at each offset.
Leer descripción completaMostrar menos
A few kilobytes of input therefore consumes seconds to minutes of single-threaded CPU, and the HANG_RE substitution performed before matching does not collapse the pattern. TweetTokenizer is intended for tokenizing untrusted social-media text, so any service that applies it, or the module-level casual_tokenize, to submitted text can be stalled per request without authentication. Version 3.10.1 bounds the label repetition.
CVSS
- Versión: 4.0
- Vector: CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N/E:X/CR:X/IR:X/AR:X/MAV:X/MAC:X/MAT:X/MPR:X/MUI:X/MVC:X/MVI:X/MVA:X/MSC:X/MSI:X/MSA:X/S:X/AU:X/R:X/V:X/RE:X/U:X
- Puntuación base: 8.7
Probabilidad de explotación (EPSS)
- Probabilidad de explotación en los próximos 30 días: 0.74%
- Percentil entre todas las CVEs puntuadas: 53
- Fecha de la puntuación: 6/10/2026
EPSS (Exploit Prediction Scoring System, de FIRST) estima la probabilidad de que una vulnerabilidad sea explotada en 30 días. Complementa a CVSS (impacto) y a CISA KEV (explotación confirmada).
🎯 Técnicas ATT&CK
Cómo se explota esta vulnerabilidad y qué consigue el atacante, en el lenguaje de MITRE ATT&CK.
- Explotación
T1190Exploit Public-Facing Applicationinitial access85 % - Impacto principal
T1499.004Application or System Exploitationimpact80 %
Vulnerabilidad de ReDoS en regex de TweetTokenizer expuesta en API de NLTK sin autenticación (AV:N/PR:N). Causa negación de servicio por consumo de CPU (VA:H indica indisponibilidad).
Inferido por nuestro agente de análisis a partir de la descripción oficial, el vector CVSS y la CWE, y comprobado por un supervisor. Puede contener errores.
🛡️ Mitigaciones ATT&CK que cubren estas técnicas
Tecnologías afectadas (1)
⚠ Inferidas por IA a partir de la descripción — NVD aún no ha analizado esta CVE; no son CPE verificados.
CWE
- CWE-1333
Referencias
- https://github.com/nltk/nltk
- https://github.com/nltk/nltk/blob/3.9.4/nltk/tokenize/casual.py
- https://github.com/nltk/nltk/issues/3704
- https://github.com/nltk/nltk/releases/tag/v3.10.1
- https://www.vulncheck.com/advisories/nltk-tweettokenizer-url-pattern-backtracks-catastrophically-on-naked-domain-like-input
- https://github.com/nltk/nltk/issues/3704
JSON original (NVD)
Mostrar
{
"id": "CVE-2026-72818",
"cveTags": [],
"metrics": {
"ssvcV203": [
{
"source": "134c704f-9b21-4f2e-91b3-4a467353bcc0",
"ssvcData": {
"id": "CVE-2026-72818",
"role": "CISA Coordinator",
"options": [
{
"exploitation": "poc"
},
{
"automatable": "yes"
},
{
"technicalImpact": "partial"
}
],
"version": "2.0.3",
"timestamp": "2026-08-21T11:05:49.886911Z"
}
}
],
"cvssMetricV31": [
{
"type": "Secondary",
"source": "disclosure@vulncheck.com",
"cvssData": {
"scope": "UNCHANGED",
"version": "3.1",
"baseScore": 7.5,
"attackVector": "NETWORK",
"baseSeverity": "HIGH",
"vectorString": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"integrityImpact": "NONE",
"userInteraction": "NONE",
"attackComplexity": "LOW",
"availabilityImpact": "HIGH",
"privilegesRequired": "NONE",
"confidentialityImpact": "NONE"
},
"impactScore": 3.6,
"exploitabilityScore": 3.9
}
],
"cvssMetricV40": [
{
"type": "Secondary",
"source": "disclosure@vulncheck.com",
"cvssData": {
"Safety": "NOT_DEFINED",
"version": "4.0",
"Recovery": "NOT_DEFINED",
"baseScore": 8.7,
"Automatable": "NOT_DEFINED",
"attackVector": "NETWORK",
"baseSeverity": "HIGH",
"valueDensity": "NOT_DEFINED",
"vectorString": "CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N/E:X/CR:X/IR:X/AR:X/MAV:X/MAC:X/MAT:X/MPR:X/MUI:X/MVC:X/MVI:X/MVA:X/MSC:X/MSI:X/MSA:X/S:X/AU:X/R:X/V:X/RE:X/U:X",
"exploitMaturity": "NOT_DEFINED",
"providerUrgency": "NOT_DEFINED",
"userInteraction": "NONE",
"attackComplexity": "LOW",
"attackRequirements": "NONE",
"privilegesRequired": "NONE",
"subIntegrityImpact": "NONE",
"vulnIntegrityImpact": "NONE",
"integrityRequirement": "NOT_DEFINED",
"modifiedAttackVector": "NOT_DEFINED",
"subAvailabilityImpact": "NONE",
"vulnAvailabilityImpact": "HIGH",
"availabilityRequirement": "NOT_DEFINED",
"modifiedUserInteraction": "NOT_DEFINED",
"modifiedAttackComplexity": "NOT_DEFINED",
"subConfidentialityImpact": "NONE",
"vulnConfidentialityImpact": "NONE",
"confidentialityRequirement": "NOT_DEFINED",
"modifiedAttackRequirements": "NOT_DEFINED",
"modifiedPrivilegesRequired": "NOT_DEFINED",
"modifiedSubIntegrityImpact": "NOT_DEFINED",
"modifiedVulnIntegrityImpact": "NOT_DEFINED",
"vulnerabilityResponseEffort": "NOT_DEFINED",
"modifiedSubAvailabilityImpact": "NOT_DEFINED",
"modifiedVulnAvailabilityImpact": "NOT_DEFINED",
"modifiedSubConfidentialityImpact": "NOT_DEFINED",
"modifiedVulnConfidentialityImpact": "NOT_DEFINED"
}
}
]
},
"affected": [
{
"source": "disclosure@vulncheck.com",
"affectedData": [
{
"repo": "https://github.com/nltk/nltk",
"vendor": "nltk",
"product": "nltk",
"versions": [
{
"status": "affected",
"version": "0",
"lessThan": "3.10.1",
"versionType": "semver"
},
{
"status": "unaffected",
"version": "3.10.1",
"versionType": "semver"
}
],
"packageURL": "pkg:pypi/nltk",
"packageName": "nltk",
"programFiles": [
"nltk/tokenize/casual.py"
],
"collectionURL": "https://pypi.org",
"defaultStatus": "unaffected"
}
]
}
],
"published": "2026-08-20T22:18:05.087",
"references": [
{
"url": "https://github.com/nltk/nltk",
"source": "disclosure@vulncheck.com"
},
{
"url": "https://github.com/nltk/nltk/blob/3.9.4/nltk/tokenize/casual.py",
"source": "disclosure@vulncheck.com"
},
{
"url": "https://github.com/nltk/nltk/issues/3704",
"source": "disclosure@vulncheck.com"
},
{
"url": "https://github.com/nltk/nltk/releases/tag/v3.10.1",
"source": "disclosure@vulncheck.com"
},
{
"url": "https://www.vulncheck.com/advisories/nltk-tweettokenizer-url-pattern-backtracks-catastrophically-on-naked-domain-like-input",
"source": "disclosure@vulncheck.com"
},
{
"url": "https://github.com/nltk/nltk/issues/3704",
"source": "134c704f-9b21-4f2e-91b3-4a467353bcc0"
}
],
"vulnStatus": "Awaiting Analysis",
"weaknesses": [
{
"type": "Secondary",
"source": "disclosure@vulncheck.com",
"description": [
{
"lang": "en",
"value": "CWE-1333"
}
]
}
],
"descriptions": [
{
"lang": "en",
"value": "The URLS regular expression in nltk/tokenize/casual.py, compiled into TweetTokenizer.WORD_RE and applied by TweetTokenizer.tokenize, contains a naked-domain branch whose domain-label prefix [a-z0-9]+(?:[.\\-][a-z0-9]+)* is unbounded. Input consisting of many alternating label separators can be partitioned in exponentially many ways, and because the branch also requires a trailing top-level domain that such input never supplies, the engine explores those partitions before failing at each offset. A few kilobytes of input therefore consumes seconds to minutes of single-threaded CPU, and the HANG_RE substitution performed before matching does not collapse the pattern. TweetTokenizer is intended for tokenizing untrusted social-media text, so any service that applies it, or the module-level casual_tokenize, to submitted text can be stalled per request without authentication. Version 3.10.1 bounds the label repetition."
}
],
"lastModified": "2026-09-24T20:02:50.260",
"sourceIdentifier": "disclosure@vulncheck.com"
}