A particularly serious security vulnerability has been identified in LMCache, the open-source software used to accelerate artificial intelligence applications and large language models (LLMs), including in vLLM. The vulnerability could allow a remote, unauthenticated attacker to execute arbitrary code on the cache server, without first needing to log in to the system.

The vulnerability has been recorded as CVE-2026-105192 and has received a CVSS score of 9.8/10, which places it in the critical category. The most significant problem for administrators is that no official version of LMCache has been released to fix the issue so far.
How the attack works
The problem specifically concerns the multiprocess functionality of LMCache. In this configuration, the cache operates as a separate server and the worker processes of an LLM communicate with it via ZeroMQ.
This communication socket does not have an authentication mechanism. This means that if the server is accessible from the network, an attacker can send a specially crafted message and cause code execution.
See also: SGLang: Vulnerability allows RCE via malicious GGUF model files
The core of the problem lies in the use of Python's pickle library for data deserialization. LMCache pickles certain data before performing the necessary checks on the content and type of the message. Thus, a maliciously crafted packet can lead to the execution of commands controlled by the sender.
The level of access an attacker gains depends on the permissions of the LMCache. JFrog points out that in the project's official container images, the process runs as root, which significantly increases the potential impact of a successful attack.

A regulation judges the exposure
However, not all LMCache installations are equally vulnerable. By default, the multiprocess server only listens on localhost, limiting access to the same machine.
The risk increases when the administrator chooses a routable IP address so that the server is available to other nodes. This is common in multi-node installations, where the cache needs to be used by different machines.
Kubernetes installations also require special attention . The deployment example provided by LMCache itself configures the server to listen on all available network interfaces . Such a configuration can significantly increase the attack surface, especially when there are no additional access restrictions.
Which versions are affected?
According to JFrog, the vulnerability affects LMCache versions from 0.3.9 to 0.5.5, while release candidate 0.5.6 and the development branch.
See also: Insecure deserialization in NLTK: CVE-2026-78683
The issue was discovered by security researcher Yuval Moravchick of JFrog and was disclosed on October 7. Until an official patch is released, administrators are advised to avoid exposing the multiprocess server to routable networks.
What should administrators do?
The basic temporary defense is to restrict network access. The LMCache server should only be available on localhost or a fully trusted internal network of the cluster. A firewall can further restrict the connections allowed, but is not a complete solution if any untrusted host can still communicate on the port.
🔒 Protect your privacy with Proton VPN
Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.
- ✔ No-logs, based in Switzerland (except 14-Eyes)
- ✔ NetShield: blocks ads, trackers & malicious domains
- ✔ Covers all devices — free version available
The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.
At the same time, it is important for administrators to avoid running such services with root privileges when not necessary. The principle of least privilege can significantly limit the consequences in the event of an exploit.
Additional references and the vLLM
The disclosure of the critical vulnerability comes shortly after six other security reports for LMCache. The reports concern, among other things, possible unauthenticated access to cache data across tenants and network services that can perform actions without login. However, they have not yet been accompanied by a CVE or official confirmation from the maintainers.

Meanwhile, a different issue in vLLM has already been fixed. CVE-2026-105756 could cause a denial-of-service via an invalid cache_salt in installations using the LMCache multiprocess connector. The issue was fixed prior to vLLM 0.30.0 and did not allow arbitrary code execution.
See also: Critical bug leaves Hugging Face's LeRobot exposed
The case highlights a broader issue in AI infrastructure security: services originally designed for internal communication between components can become critical entry points when exposed to the network. This vulnerability is also reminiscent of the ShadowMQthat emerged in other AI inference frameworks in 2025.
Until an official fix is available, limiting LMCache's network exposure is the most important safeguard. Especially in production infrastructures, Kubernetes clusters, and multi-node environments, this setting should be checked immediately.
