Summary A frontend-legal multi-request speculative workload can make vLLM produce an out-of-vocabulary recovered token equal to vocab_size, convert that value to -1 when choosing the next live token for a request, and t…
| CVE ID | CVE-2026-54234 |
| Vendor | Unknown Vendor |
| Affected Product | Unknown Product |
| Vulnerability Type | Security Vulnerability |
| CVSS Score | 7.5 (HIGH) |
| EPSS Score | 0.3% probability of exploitation in the next 30 days |
| Actively Exploited | ❌ No known exploitation |
| Patch Status | Pending Vendor Disclosure |
| Reported By | CYBERDUDEBIVASH SENTINEL APEX Intelligence (via sentinel_apex) |
#
A frontend-legal multi-request speculative workload can make vLLM produce an out-of-vocabulary recovered token equal to vocab_size, convert that value to -1 when choosing the next live token for a request, and then feed that -1 back into the next drafter input ids. On Qwen3 GPTQ this reaches the worker-side drafting / attention path and crashes the engine with a GPU device-side assert.
The same issue is reachable through the public gRPC request surface by sending a specific overlapping Generate / Abort sequence.
#
shared vLLM engine worker requests from completing until the worker is restarted -
Sigma rules, YARA signatures, IOC table, and SIEM queries for Splunk, Elastic, Sentinel, and Chronicle — deployable in 5 minutes.