top of page

LMCache CVE-2026-105192: One Message to Port 5555 Runs Code as Root on the Box Holding Your Model's Memory. There Is No Patch, and Our Edge Can't See It.

Writer: Patrick Duggan
Patrick Duggan
7 hours ago
5 min read

LMCache is the piece of an AI serving stack most people never look at. It sits next to vLLM and stores the model's working memory, the key-value cache it builds while reading a prompt, so the next request with the same context does not have to compute it again. Big deployments share that cache across machines, because that is where the savings are. On October 7, JFrog showed that sharing it this way hands remote code execution to anyone who can reach one network port. The flaw is CVE-2026-105192, it scores 9.8, and as of JFrog's write-up there is no fixed version and no advisory from the project.



The bug


Credit for the research belongs to Yuval Moravchick of JFrog's security research team, summarized by The Hacker News and cataloged at Strix.


In its multiprocess mode, LMCache runs a server that listens on a ZeroMQ socket, port 5555 by default, so worker processes can register and share cache blocks. Messages arrive packed with msgpack. One msgpack extension type, code 1, hands its contents to Python's pickle.loads. Pickle is not a data format so much as a small program format: unpacking it can run code. And this unpacking happens while the request's arguments are being decoded, before LMCache has checked what kind of message it is, and with no authentication at all.


So one message is enough. Send a single crafted ZeroMQ message to the port and you run code as whatever user LMCache runs as. In the project's official container images, that user is root.


Affected versions are 0.3.9 through 0.5.5, plus the 0.5.6 release candidates and the development branch.





Safe by default, unsafe when used as intended


By default the server listens only on the local machine, so another host cannot reach it. That sounds like the end of the story. It is not, because the reason to run this server at all is to share the cache across machines, and to do that you give it an address other machines can reach. JFrog points out that the project's own Kubernetes example does exactly that. When the documented way to deploy something is the vulnerable way, "off by default" protects only the people who were not really using it.




Reachable does not have to mean the internet. Inside a GPU cluster, every pod or container on the same network can usually reach every other one. One compromised notebook, one poisoned dependency in a sidecar, or one tenant on a shared cluster is enough.



What a root shell on that box is worth


This is not a box that just does math. The cache it holds is built from your users' prompts and your system prompts, so whoever owns it is sitting next to recent conversations, retrieved documents and agent context. The node usually carries credentials for model storage and cloud APIs. And it has GPUs, which is the part criminals have been monetizing all year: we have written about crews mining cryptocurrency on hijacked AI servers, from a Monero miner riding a Langflow bug to NadMesh, a botnet built to hunt AI tools. An unauthenticated root bug on a GPU node is exactly what those crews go looking for. JFrog did not report any exploitation, and neither do we.



What our edge can and cannot tell you


We block attacks at the edge of our own sites and record what we see there. For a bug like the Atlassian one this week, that gives us a real measurement. For this one it gives us nothing, and we would rather say so than imply coverage. Our edge only sees HTTP and HTTPS traffic to our own domains. A ZeroMQ message to port 5555 on someone's GPU cluster never passes through it, so "we have seen no attempts" would be true and meaningless. We hold no indicators for this campaign because there is no campaign on record yet, only a bug.


The measurement that matters is yours: whether anything in your environment is listening on that port where it should not be.



We measured the internet side. That number is the trap.


Update, October 8: we went looking for the exposed surface with Shodan, counting only. We did not probe anyone.


  • Anything answering on port 5555: 969,065 hosts, almost all of it web servers, Android debug bridges and cameras.

  • Services Shodan identifies as ZeroMQ, on any port: 8,201.

  • ZeroMQ services answering on port 5555: 2.

  • Pages that mention "lmcache" at all: 6.

  • ZeroMQ by GPU-focused neocloud: Nebius 2, Denvr 1, and zero at CoreWeave, Lambda, Crusoe, RunPod, Together and Fluidstack.

  • ZeroMQ by large and sovereign clouds: Korea Telecom 950, Google 173, OVH 164, Microsoft 133, Tencent 123, Huawei Cloud 110, e& 73, Aliyun 62, Orange 43, Deutsche Telekom 9, NAVER 2, STACKIT 0.

Read those carefully. ZeroMQ is not LMCache: plenty of other software speaks it (ports 4505 and 4506 are SaltStack, for example), so none of the per-provider numbers says anyone there runs LMCache. And Shodan mostly probes port 5555 as a web or Android port, so 2 is a floor, not a census.


The conclusion is still sharp. On the open internet this bug is close to invisible. That is the dangerous part. The exposure lives east-west, inside GPU clusters, on the pod network, where the project's own Kubernetes example puts the server on a routable address. No internet scanner will ever count it, the cloud provider's edge never sees it, and the tenant who launched a one-click vLLM template assumes the provider handled it. The control holds on both sides and the handoff in the middle belongs to nobody. If you run managed vLLM templates or GPU Kubernetes for customers, check your templates and default network policies for port 5555, and tell your tenants. Do not wait for a scanner to tell you, because it will say 2.



What to do while there is no patch


Find every LMCache deployment and check what address its multiprocess server binds to. If it is anything other than the local machine, treat it as exposed.


Restrict port 5555 so only the specific worker hosts that need it can reach it, using Kubernetes network policies or host firewalls. JFrog is clear that this reduces the risk and does not remove it: anything that can reach the port can still own the box.


Run LMCache as a non-root user if your deployment allows it, so a successful exploit gets less.


If you do not need a cache shared across machines, do not run it that way.


Watch the project for a release that stops passing network input to pickle, and upgrade when it lands. Until then, this is a 9.8 with no patch in a component that sits next to your users' prompts.


Was this useful? Rate this post. The widget is at the bottom of the page, and we read every response.




How do AI models see YOUR brand?

AIPM has audited 250+ domains. 15 seconds. Free while still in beta.



Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=lmcache-cve-2026-105192-unpatched-zeromq-pickle-rce-vllm-kv-cache



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page