A practical architecture for remote KV caching with Valkey
Reuse KV cache across inference workers with a shared Valkey cache, RDMA/EFA networking, plus NVMe and object storage tiers, all powered by Momento.

An inference fleet can process the same context many times even when prefix caching is enabled. A worker keeps useful state locally, but the next request lands elsewhere. Long document prefixes get recomputed while their earlier results sit on another machine or have already been evicted.
A shared Valkey tier gives those results a place that every compatible worker can reach. Momento’s RDMA and AWS EFA modules provide a direct path between inference GPUs and KV blocks in Valkey RAM. Its NVMe module adds local disk capacity, and its object storage connector asynchronously flushes data to inexpensive long term storage. These capabilities let you retain reusable state beyond the lifetime and memory budget of an individual worker.
Consider an assistant answering questions about an operations manual. Each prompt begins with the same system instructions and manual, followed by a different question. The objective is to reuse that shared prefix wherever the next question runs.
Keep the useful state
KV cache contains the attention keys and values computed for prior tokens. Prefill processes the input and establishes this state. Decode uses it to generate the response and extends it as new tokens arrive. Keeping those tensors avoids recreating them during generation.
Prefix caching extends reuse across requests. If a later prompt has the same token prefix and compatible model state, the engine can reuse completed cached blocks. It still processes the new question and generates an answer. Matching subject matter is insufficient. Changing the instructions or moving the manual later in the prompt can change the reusable prefix.
Remote reuse also differs from prefill and decode disaggregation. Disaggregation runs the two phases on separate worker pools and requires a state handoff between them. A shared cache can support that design, but it also benefits workers that run both phases together. The manual example uses such workers.
Put Valkey between inference and storage
Deploy Valkey on storage nodes separate from the inference workers. GPU memory holds state needed for active inference. Valkey RAM holds reusable blocks shared across the fleet. Local NVMe on the Valkey nodes and an object bucket extend the retained working set.
Separating storage from compute lets you add cache capacity without buying another inference GPU. It also gives new or replaced workers access to previously computed prefixes. Valkey supplies an open source storage foundation with a module extension mechanism. Momento’s modules add the transfer and storage capabilities needed for this architecture.
The following reference design stages lower tier blocks through Valkey RAM. Placement, promotion, and coordination are architectural choices for this design.

Direct GPU transfers use Valkey RAM. Lower tier blocks return through RAM in this reference design. Object writes run asynchronously, separately from immediate reuse.
The runtime integration identifies reusable blocks, coordinates availability, and arranges for tensors to become usable GPU cache state. Lookup and bulk transfer have separate roles. Finding a block does not mean that the GPU is ready to consume it.
Suppose worker A has processed the manual and exported eligible completed blocks. Worker B receives another question and checks its local cache first. For a local miss, the integration looks for compatible prefix blocks in Valkey. A RAM hit uses Momento’s RDMA or EFA module to move the blocks into worker B’s GPU memory. The runtime then processes the unmatched prompt suffix and continues generation.
If no useful blocks are available, worker B computes the prefix. Its new blocks can be admitted to shared storage for later requests. Keep retrieval bounded so a slow lookup does not turn a recoverable miss into a long wait. Once RAM reuse works, the next step is retaining more useful prefixes than RAM alone can hold.
Extend the working set with NVMe and object storage
Give each tier a distinct role. Valkey RAM serves shared prefixes used frequently. Momento’s NVMe module provides more capacity on the storage node for blocks worth keeping outside RAM. Object storage provides a larger retention tier for contexts that may return after hours, sessions, or worker replacements.
In this reference design, requested NVMe blocks are staged into Valkey RAM and then transferred to the GPU through the same fast path. The disk is local to Valkey, so inference workers access it through the storage service. They do not each need a disk copy of every manual.
Momento’s object connector flushes data asynchronously. A block already available in RAM can be reused while its object copy is being written. That decouples immediate reuse from the upload, but it does not make every cache write immediately durable. If the only usable copy disappears before the flush completes, the prefix may need recomputation.
For an object hit, the proposed return path reads retained blocks into Valkey RAM before loading the GPU. Recoverable lookup metadata must accompany retained state so the system can identify compatible objects after a restart. Object storage extends the reuse window, while retrieval still has to fit the request’s latency budget. Choose a storage class that supports the access times you need.
Admission and retention policies keep this hierarchy useful. Prefer blocks likely to be reused over a stream of one-time prompts. Keep copies available until transfers using them finish. Coordinate expiry across lookup metadata and lower tiers so deleted or incompatible state cannot appear as a usable hit. Evaluate the tradeoffs between capacity and coordination when defining your architecture.
Workers still perform attention using state loaded into their GPU cache. The remote tiers preserve reusable blocks between uses. If reading and installing an older prefix costs more than recomputing it, a bounded fallback is the better serving decision. That comparison makes the network part of the cache design.
Match the network to the workload
Measure effective block transfer bandwidth under serving load, including lookup, lower tier reads, and GPU installation. Network line rate alone cannot tell you whether a hit will beat prefill. Concurrent inference traffic and cache fills also compete for bandwidth.
RDMA reduces overhead by moving data between registered memory regions without the usual socket data path. With compatible hardware and software, GPU direct transfer avoids staging the payload through inference host memory. Momento supplies the Valkey side of that direct path.
On AWS, use Momento’s EFA module with compatible GPU and storage endpoints. Confirm the instance pairing and GPU direct support before provisioning. EFA device traffic stays within an Availability Zone and VPC. Place the object bucket in the same Region and access it through normal object APIs.
Connect the serving stack
Momento Cache is the leading provider of Valkey, an open-source fork of Redis built for real-time workloads. With Momento Cache and Valkey, your team can enjoy the same high-performance infrastructure that powers global leaders like Snap, FOX, Coinbase, and Capcom.
Momento provides direct integrations with SGLang, vLLM, and LMCache. Choose the path that fits your serving stack. For the manual example, start with the direct vLLM integration if vLLM already runs your workers. If LMCache already manages offloading, use Momento’s LMCache integration. The direct SGLang integration offers the corresponding entry point for SGLang deployments.
Both workers need the same model revision and adapters, matching tokenized prefixes, and compatible tensor representation and parallel layout. Give a new model or cache format its own namespace. Scope sharing to authorized tenants or trust groups, and protect storage access independently of prefix matching. A content hash alone is not an authorization boundary. These requirements belong in the integration and deployment configuration so application code can stay focused on requests.
How to measure the impact of a shared cache
A useful first deployment needs two compatible inference workers, one separate Valkey node with RAM and NVMe, and an object bucket. Obtain the Momento modules and runtime integration for the chosen topology. Start by proving shared RAM reuse before introducing disk and object retrieval.
Send a long manual prefix to worker A, then send a different question with that same prefix to worker B. Keep local prefix caching enabled in your baseline. Confirm that B lacks a local copy and obtains compatible blocks from shared storage. Compare its time to first token with recomputing the prefix under the same serving conditions.
Next, grow the working set beyond the RAM tier to exercise NVMe. Test object retrieval using blocks whose flush has completed and whose faster copies have been removed through supported controls. Restart an inference worker to check that reuse survives worker turnover. Test an unavailable remote tier and verify bounded retrieval followed by recomputation.
Measure tail latency and transfer time as concurrency rises. Track hits and retrieved bytes by tier, GPU memory pressure, and the object flush backlog. Include cache memory, disk, object requests, and network costs when judging the result. A hit rate without its retrieval cost can make an expensive cache look successful.
Accelerate your KV caching with Valkey and Momento
Valkey is a practical solution when repeated prefixes outlive one worker’s cache and loading their state beats recomputation. Momento’s modules connect that shared RAM tier to inference GPUs and extend it onto disk and object storage. Looking for advice or help to trying it out? Get in touch with Momento and one of our solution engineers will be happy to take a look!