<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"
  xmlns:content="http://purl.org/rss/1.0/modules/content/"
  xmlns:dc="http://purl.org/dc/elements/1.1/"
  xmlns:atom="http://www.w3.org/2005/Atom"
  xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
>
<channel>
  <title>Momento Japan Feed</title>
  <atom:link href="https://www.gomomento.com/jp/feed" rel="self" type="application/rss+xml" />
  <link>https://www.gomomento.com/jp/</link>
  <description>Momentoの日本向け更新情報。</description>
  <lastBuildDate>Wed, 05 Aug 2026 09:00:00 GMT</lastBuildDate>
  <language>ja</language>
  <sy:updatePeriod>hourly</sy:updatePeriod>
  <sy:updateFrequency>1</sy:updateFrequency>
  <image>
    <url>https://www.gomomento.com/wp-content/uploads/2024/06/cropped-favicon-green-32x32.png</url>
    <title>Momento Japan Feed</title>
    <link>https://www.gomomento.com/jp/</link>
    <width>32</width>
    <height>32</height>
  </image>
  <item>
    <title>BuffConf 2026: Signal Over Noise, One Year Later</title>
    <link>https://www.gomomento.com/jp/blog/buffconf-2026-signal-over-noise-one-year-later/</link>
  <dc:creator><![CDATA[Lionel Bringuier]]></dc:creator>
    <pubDate>Wed, 05 Aug 2026 09:00:00 GMT</pubDate>
  <category><![CDATA[Events]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/buffconf-2026-signal-over-noise-one-year-later/</guid>
    <description><![CDATA[<p>A recap of two days of honest engineering conversations, practical lessons, and the community behind BuffConf 2026.</p>]]></description>
    <content:encoded><![CDATA[<p><img src="https://www.gomomento.com/assets/content/blog/buffconf-2026-signal-over-noise-one-year-later/recap-blog-banner.png" alt="Audience members seated in a theater for BuffConf 2026."></p>
<p>Last year, <a href="https://buffconf.gomomento.com/past-sessions">the first edition of Buffer-Free Video)^AI</a> left me in disbelief that a scrappy ten-week sprint had pulled 150 streaming professionals into one room. This year the question in my head was different. Not “can we do it again?” but “can we keep the signal-to-noise ratio that made the first one special?”</p>
<p>I believe we did, and the numbers back it up.</p>
<h2 id="the-room-we-wanted">The Room We Wanted</h2>
<p>Here is the stat I am proudest of: we had 117 attendees show up out of 124 registrations. That is a 94.4% show-rate. If you have ever organized a tech conference, you know the typical number hovers around 60%. People register, life happens, half the badges never get picked up.</p>
<p>That did not happen here, and it tells you something about who was in the room. Definitely not people collecting free T-shirts, stickers, Mo plushies or vintage green Elemental swag (thanks Kiran!). Of the seven no-shows, six were from sponsoring organizations who had other folks covering the event.</p>
<p>Those 117 people represented 41 unique organizations, from the largest broadcasters, tech platforms, video infrastructure providers… the full spectrum of who actually builds and runs media workloads at scale. When you get that kind of density of real operators in one place, the hallway conversations become as valuable as the sessions. Actually, maybe even more.</p>
<p><img src="https://www.gomomento.com/assets/content/blog/buffconf-2026-signal-over-noise-one-year-later/in-the-room-buffconf-2026.png" alt="The BuffConf 2026 audience."></p>
<h2 id="we-pre-screened-the-pitches-out">We Pre-Screened the Pitches Out</h2>
<p>Here is the thing we are strict about. BuffConf is not a place to sell. Every proposed session got human-screened, and anything that read like a vendor pitch got sent back. We prioritized end-user and practitioner voices, the people who have actually been paged at 3am when things flinched.</p>
<p>The result was 22 technical sessions with 29 speakers from 17 companies. And critically, 40% of our speakers were non-vendors. We are talking real production stories from operators like Paramount, Netflix, FOX, CBS Interactive, Meta and Red Bull Media House. Not “here is our roadmap.” More like “here is what we tried, what failed and what finally held up.”</p>
<p>It would be easy to fill an agenda with vendors who have marketing budgets and room for another video nerd conference (all the more when the traditional bigger one is seemingly canceled silently this year… of course you know what I’m talking about!). It is much harder, and also much more valuable, to get the operators to stand up and tell the unvarnished truth about their infrastructure. When they do, everyone in the room learns something they can use in their own work.</p>
<h2 id="what-the-presentations-covered">What The Presentations Covered</h2>
<p>Three main themes ran through the two days, and a handful of sessions captured each one.</p>
<p><strong>Production workflows at scale.</strong> I opened with “104 Matches. Five Billion Viewers. Zero Margin for Error,” walking through the FIFA 2026 World Cup numbers and what global live scale actually demands, with no excuses. Spencer Shanson from Paramount then told the story of scaling a FAST platform from a handful of channels to thousands. That kind of journey, from a few streams to an entire catalog of linear channels, is exactly the operational reality most of the room is living on a daily basis, and hearing it from someone who has actually done it is worth more than any architecture diagram. And when “production at scale” means “at Netflix’s scale”, who better than Sunjeet Singh and William Schor could explain how Netflix handles over one billion RPS on 120,000 nodes?</p>
<p><img src="https://www.gomomento.com/assets/content/blog/buffconf-2026-signal-over-noise-one-year-later/spencer-shanson-buffconf-2026.png" alt="Spencer Shanson at BuffConf 2026."></p>
<p><strong>Practical AI, and honest boundaries around it.</strong> This was the part I found most refreshing. Nobody stood up and told us AI was going to solve everything. It’s still a journey, and our keynote speaker, Manish Rao, showed us all the progress that AWS Elemental Inference has made since NAB. It’s staggering how three months can mean a world of difference at the current speed of AI innovation. Edwin Rivera from CBS Interactive showed how CBS Sports built an AI-powered “For You Page,” blending editorial curation with embedding-based recommendations to actually move engagement. Content discovery in sports is one of the richest problem spaces in our industry right now. And Jeffrey Kember from NVIDIA walked through what accelerated computing is doing for encoding, quality, and real-time media workflows. A reminder that we are genuinely at the beginning of AI for media, not the end. Across the AI sessions, we also saw token efficiency work borrowed straight from video compression concepts, automated QC loops, and an edge anti-piracy simulation harness. Real implementations, with the scars to prove it, paired with hands-on framework training so engineering teams could take something home and build.</p>
<p><img src="https://www.gomomento.com/assets/content/blog/buffconf-2026-signal-over-noise-one-year-later/edwin-rivera-buffconf-2026.png" alt="Edwin Rivera at BuffConf 2026."></p>
<p><strong>The client-side reality.</strong> Our panel on Optimizing Playback SDKs and Device Capabilities brought together Jonathan Colwell from FOX, Joshua Lamb from Red Bull Media House, and David Van der Voort “DVD” from Paramount. Three very different takes on the same daily chaos: supporting every device under the sun, with every usage pattern. That fragmentation is a grind everyone in playback knows intimately, and getting three operators to compare notes in the open was exactly the kind of session you cannot get anywhere else.</p>
<p><img src="https://www.gomomento.com/assets/content/blog/buffconf-2026-signal-over-noise-one-year-later/panel-buffconf-2026.png" alt="The BuffConf 2026 playback panel."></p>
<p>We kept the format balanced through all of it. Deep technical talks, real networking time, and interactive workshops where people could get their hands dirty. Room to think, argue, and go deeper.</p>
<h2 id="the-qa-is-still-the-best-part">The Q&#x26;A Is Still the Best Part</h2>
<p>Same as last year, the Q&#x26;A sessions were electric. A technical crowd asked technical questions, and every speaker rose to meet them. Those exchanges are where a prepared talk turns into something better, where you get past the slides into the messy details that only come out when a peer in the audience pushes on exactly the right point.</p>
<p>That is the whole reason we do this in person. You cannot replicate that online.</p>
<h2 id="thank-you">Thank You</h2>
<p>A conference only works because of the people who believe in it before it exists.</p>
<p>To our speakers, from Paramount, Netflix, CBS Interactive, FOX, Red Bull Media House, NVIDIA, Meta, AWS, Visionular, Bitmovin and every company that sent someone to share hard-won lessons: thank you for choosing substance over spin. You are the reason people fly in.</p>
<p>To JP Saibene and Nicolas Gonzalez from Uruguay, to Ali Begen from Tűrkiye and to everyone who traveled a long way to be in Seattle: thank you for bringing the curiosity and the generosity that makes this community what it is.</p>
<p>To the Momento crew who pulled this together, and in particular Hannah, Mike, Kelsey, Maéline and everyone who spent months making a two-day event look effortless: thank you, I’m incredibly grateful and lucky to be working with you all.</p>
<p>And to the people who showed up: this community was the event. Your questions, your stories, your willingness to say “actually, that broke for us too” is what makes this community worth gathering. The streaming landscape keeps moving fast, especially where AI meets live video. We were all here to discuss the future of video, one honest conversation at a time.</p>
<p><a href="https://buffconf.gomomento.com/">See you next year.</a></p>]]></content:encoded>
  </item>
  <item>
    <title>The most expensive word in inference is &quot;now&quot;</title>
    <link>https://www.gomomento.com/jp/blog/the-most-expensive-word-in-inference-is-now/</link>
  <dc:creator><![CDATA[Allen Helton]]></dc:creator>
    <pubDate>Thu, 30 Jul 2026 09:00:00 GMT</pubDate>
  <category><![CDATA[AI/ML]]></category>
  <category><![CDATA[Performance]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/the-most-expensive-word-in-inference-is-now/</guid>
    <description><![CDATA[<p>Async agents don't have a human staring at the screen, and that changes the economics of inference more than you'd think. Practical takeaways after watching Meryem Arik's talk on inference for async agents.</p>]]></description>
    <content:encoded><![CDATA[<p>A few weeks ago I handed off a pretty hefty refactor to Claude, closed my laptop, and went out to do some farm chores. I had no idea how long the job was going to take, and I didn’t care. I walked back in a couple hours later and it was done.</p>
<p>That should feel stranger than it does. When we typically think about inference, speed is a whole thing. Time to first token (TTFT), tail latency, and shaving milliseconds off the round trip so a human can stay in the flow are all you see in articles these days. Yet there I was running one of the most token-hungry things from my laptop, and I didn’t really care how long it took.</p>
<p><a href="https://www.linkedin.com/in/meryemarik/">Meryem Arik</a>, CEO of <a href="https://doubleword.ai/">Doubleword</a>, was able to articulate really well how I felt about it in her recent talk <a href="https://youtu.be/M4J940vgjgI"><em>Inference for Async Agents in Production</em></a>. She posed the simple question “<em>who’s waiting?</em>” that made it all click for me.</p>
<p>When nobody’s waiting, the cost of inference looks entirely different.</p>
<h2 id="schedule-inference-like-nobodys-watching">Schedule inference like nobody’s watching</h2>
<p>In a synchronous system like chat, a human is directly in the flow. Make them wait without any feedback and they start clicking around wondering if anything is happening. In an async system, an agent might fan out, run for thirty minutes, and ping you when it’s done. Nobody is watching the screen. In many instances, it doesn’t really matter whether it comes back in 35 minutes or 45, because the user has gone off to do something else.</p>
<p>To me, this reads like a UX detail, but it drives almost every other decision in your stack.</p>
<p>“Who’s waiting?” is the right place to start. More specifically, <em>how long is the work allowed to wait and what is blocked behind it?</em> Even without a person staring at the screen, a call on an agent’s critical path still has a deadline. Async workflows just give us more room to decide what that deadline actually is.</p>
<p>A human waiting is the most expensive clock in the system. “Fast enough for a person” forces things like how you provision, how much you batch, which chips you buy, and how much idle capacity you keep on standby. Saying “I want an answer now” makes everything expensive. But if you take the human out of the loop, that clock loosens its grip, along with the price tag it was setting for everything else.</p>
<h2 id="we-reinvented-batch-processing">We reinvented batch processing</h2>
<p>We’ve done this a time or two in the past with things like interactive database queries versus overnight analytics jobs. We prioritize the web request a user is staring at versus the batch job that runs at 2 a.m. The tech industry has been separating “someone is waiting” from “nobody is waiting” since before I was born. 😅</p>
<p>Inference is speed-running what databases and job queues learned decades ago, but in a good way. When a young field’s hardest problem is something the infrastructure world already recognizes, you can jump straight into borrowing patterns that are proven to work.</p>
<p>Meryem frames the core tension as the old trade-off triangle: latency, quality, throughput. You get to pick two of them.</p>
<p><img src="https://www.gomomento.com/blog/2026-07-30_the-most-expensive-word-in-inference-is-now/tradeoff-triangle.avif" alt="Tradeoff triangle between quality, low latency, and high throughput"></p>
<p>Chat apps pick latency and quality. Async agents pick quality and throughput, and they let latency slide. A long-running agent mostly needs the smartest model it can get so it stays coherent across a two-hour piece of work while it manages subagents and juggles tasks. Delivering tokens at conversational speed doesn’t buy it anything.</p>
<h3 id="batching-is-really-a-scheduling-problem">Batching is really a scheduling problem</h3>
<p>Batch-size math can be tricky. In one example from Meryem’s talk, a model running at batch size one and tuned for maximum interactivity costs roughly $4.65 per million tokens. Push enough concurrent work through the same model and tune for throughput instead, and that cost comes out closer to nine cents per million tokens.</p>
<p>In that example, the same weights and quality produce a 50x difference in cost. That makes a substantial difference to your unit economics.</p>
<p>But there’s a catch. Bigger batches produce their best economics only when you have enough compatible work waiting to keep them full. An empty slot in a batch is silicon you’re paying for that isn’t producing tokens. The hard part is keeping batches full while demand stays bursty and wildly unpredictable.</p>
<p>I’ve made this argument from the other direction before, that <a href="https://www.gomomento.com/blog/gpus-are-the-most-expensive-resource-in-tech/">GPUs are the most expensive resource in tech and we use them badly</a>. Batch size is the same thing. High utilization is paramount, and you win or lose it in the scheduler.</p>
<h3 id="the-queue-is-the-product">The queue is the product</h3>
<p>I really liked how Meryem put it: most providers lose money on real-time serverless endpoints. It comes back to the “now” tax. To promise a human fast responses against lumpy traffic, you have to provision for the peak and then eat the idle troughs in between. That idle compute is money on fire.</p>
<p>If the work has room to wait, the batch queue becomes a powerful tool. You can push work into the low periods and backfill the troughs that are already being paid for. Spot instances become fair game for soaking up spare compute when the work can tolerate interruption or be checkpointed. SLA-aware routing lets you jump an urgent request ahead of the rest. You can even reorder the queue so requests that hit the same mixture-of-experts weights run back to back and skip reloading those experts every time.</p>
<p>Once you stop committing to “now”, the scheduler/router/whatever is holding the queue becomes the center of the system. It’s  where the value collects, and it’s where everything is heading. As hardware and models commoditize, the advantage moves up into routing, placement, and orchestration, and whoever owns the “what runs next” decision owns the economics.</p>
<h2 id="start-labeling-the-wait">Start labeling the wait</h2>
<p>I mark everything real-time out of habit. Every call gets treated as if a human is holding their breath on the other end, even when the “human” is an agent off doing six other things. That results in an unnecessary “now” tax on work that doesn’t need it.</p>
<p>Doubleword has categorized latency into real-time, async (about a minute), and batch (24 hours) buckets. The 50× cost gap we mentioned earlier comes down to which bucket you pull from, and many of us (myself included) reach straight for the expensive one by reflex.</p>
<p>One of the most boring-yet-useful things you can do immediately after reading this article is going through every call your agent makes one at a time and ask whether a human is waiting on each one. You might be surprised how few are. The background call fetching context isn’t waiting. The sub-agent grinding through step four of eleven isn’t either. The nightly summarization job definitely isn’t.</p>
<p>In the near future, we’ll have routers smart enough to infer a request’s latency tier from its context, dependencies, and deadline, so we don’t have to label everything by hand. Once we figure out how to do that efficiently, system costs drop dramatically because the scheduler can optimize around how long each piece of work is actually allowed to wait. Until then, start by giving every model call a realistic latency budget instead of letting the entire workload take on “now” by default.</p>
<p>Everything we’re talking about has been around in the infrastructure world for years. The same instinct drives shared caching at Momento: preserve expensive state, reuse work you have already paid for, and keep scarce resources productive instead of idle. Inference gives those old ideas a very large new price tag.</p>
<p>If you’re wrestling with any of this, we’d love to <a href="https://www.gomomento.com/contact-us/">compare notes</a>.</p>
<p>Happy coding!</p>]]></content:encoded>
  </item>
  <item>
    <title>Consistency compounds: Valkey's journey to 200 Gbps</title>
    <link>https://www.gomomento.com/jp/blog/consistency-compounds-valkeys-journey-to-200-gbps/</link>
  <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
    <pubDate>Thu, 23 Jul 2026 16:00:00 GMT</pubDate>
  <category><![CDATA[Caching]]></category>
  <category><![CDATA[Valkey]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/consistency-compounds-valkeys-journey-to-200-gbps/</guid>
    <description><![CDATA[<p>Across four releases, I/O-threading changes removed a serial copy bottleneck, cut p99 latency, and brought large GETs to line rate on our 200 Gbps test rig.</p>]]></description>
    <content:encoded><![CDATA[<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/read-bandwidth.svg" alt="Valkey GET bandwidth by value size across versions 7.2, 8.1, 9.0, and 9.1, with the 200 Gbps NIC limit marked"></p>
<p>If you came here looking for a post from me on consistency models, I have to disappoint you. Today, I want to talk about the value of consistency in life. In life, effort is additive, but consistency is multiplicative.</p>
<p>Over the last few years, I have had the honor of watching the Valkey project blossom from an idea into an inspirational, community-driven effort with consistent improvements in each release. These improvements range from memory efficiency to availability at scale to substantial performance gains. These small improvements compound.</p>
<p>This compounding effect and the power of a driven community is perhaps best illustrated by the journey of the I/O-threading architecture and its impact on large objects from 1 MB to 64 MB in Valkey. Larger items are becoming increasingly important for inference KV caches, 4K video streaming, and other enticing use cases. At the very least, the impact of large objects should not be overlooked, as they can <a href="https://www.gomomento.com/blog/large-objects-in-valkey-9-0/">impact everyone else’s latency</a>; this post measures how fast the large objects themselves go.</p>
<p>Buckle up, because it’s about to get spicy.</p>
<h2 id="a-brief-history-of-io-threads">A brief history of I/O threads</h2>
<p>Valkey 7.2 inherited an I/O-thread model from a software stack doing its best to stay single-threaded. The main thread and I/O threads worked in coordinated phases, separated by synchronization barriers, and the <a href="https://github.com/valkey-io/valkey/blob/7.2/valkey.conf">7.2 config file</a> is candid about the result: “Usually threading reads doesn’t help much.”</p>
<p>Valkey 8.0 delivered a fundamental rearchitecture of I/O threads, <a href="https://valkey.io/blog/unlock-one-million-rps/">tripling throughput to over a million requests per second</a>. This idea was not new. In August 2023, seven months before the fork, Dan Touitou filed <a href="https://github.com/redis/redis/issues/12489">redis#12489</a>, laying out exactly this design in detail, benchmarks included.</p>
<blockquote>
<p>Redis let #12489 sit. The issue is still open in the tracker today, unassigned and without a milestone.</p>
</blockquote>
<p>Within days of the fork, the Valkey community copied the proposal verbatim into <a href="https://github.com/valkey-io/valkey/issues/22">issue #22</a>, greeted it as “a true gem,” and shipped it. The new architecture enabled continuously running I/O threads connected by queues, so reads, parses, and writes proceed on separate cores while the main thread executes commands. Valkey 8.1 delivered TLS handshake offload to I/O threads in <a href="https://github.com/valkey-io/valkey/pull/1338">#1338</a>.</p>
<p>Redis shipped a <a href="https://redis.io/blog/redis-8-0-m03-is-out-even-more-performance-new-features/">strikingly similar asynchronous I/O-threading model</a> in Redis 8 in May 2025, a year after the fork and 21 months after the design landed in its own tracker. Around the same time, we put <a href="https://www.gomomento.com/blog/valkey-turns-one-how-the-community-fork-left-redis-in-the-dust/">Valkey 8.1 and Redis 8.0 head to head</a> on small objects. Valkey 8.1 outran Redis 8.0 by 37% on writes and 16% on reads.</p>
<p>Valkey 9.0 brought <a href="https://github.com/valkey-io/valkey/pull/2078">reply copy avoidance</a>, changing the performance of large items entirely. Before this change, the main thread copied the entire object into a connection reply buffer before moving to the next command. While a large item is being copied, the entire pipeline stalls. Small objects are not serviced until the copy finishes.</p>
<p>Valkey 9.0 instead passes a reference to the I/O threads and keeps the object alive with a reference count. The I/O worker for that connection hands the object’s memory directly to <code>writev()</code>. This shortens the handoff to the I/O threads and gets the main thread back to handling requests.</p>
<p>Valkey 9.1 then redesigned communication between the main thread and I/O threads around <a href="https://github.com/valkey-io/valkey/pull/3324">lock-free queues</a>, credited in the release notes with an 8-17% throughput gain.</p>
<p>Based on these changes, we expected 8.0 to lift reads and writes for smaller items but still struggle to fill the network link for larger items. We expected 9.0 to bring GETs to line rate and 9.1 to improve writes.</p>
<p>This is what open-source competition buys everyone, including teams that never leave Redis. A performance design that sat for seven months as an unassigned issue became table stakes for both projects within two years of being filed.</p>
<h2 id="show-me-the-numbers">Show me the numbers</h2>
<p>We swept values from 1 MB to 64 MB across four Valkey releases on two nodes with 200 Gbps of bandwidth between them. The full setup is below.</p>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/8mb-release-journey.svg" alt="GET and SET bandwidth for 8 MB values across Valkey 7.2, 8.1, 9.0, and 9.1"></p>
<h3 id="valkey-72-reads-capped-at-35-gbps-writes-at-55">Valkey 7.2: reads capped at 35 Gbps, writes at 55</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-7-2-bandwidth.svg" alt="GET and SET bandwidth by value size for Valkey 7.2.13, with GET shown as a solid line and SET shown as a dashed line"></p>
<p>An 8 MB value delivers 21 Gbps on reads and 23 Gbps on writes, with p99 latencies of 174 and 164 ms. Turning on <code>io-threads-do-reads</code> moved only the 1 MB read cell.</p>
<h3 id="valkey-80-and-81-writes-take-off-large-reads-stay-put">Valkey 8.0 and 8.1: writes take off, large reads stay put</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-8-1-bandwidth.svg" alt="GET and SET bandwidth by value size, comparing Valkey 7.2.13 in red with Valkey 8.1.8 in orange"></p>
<p>The threading rebuild lifts our 8 MB write from 23 to 137 Gbps, six times faster, and 1 MB reads reach 183 Gbps. Larger reads settle at 30-33 Gbps whether the value is 8 MB or 64 MB. That flat floor points to a serial, per-byte bottleneck. Valkey 8.0.9 and 8.1.8 measured the same at every size, so one line carries both.</p>
<h3 id="valkey-90-large-gets-jump-to-line-rate">Valkey 9.0: large GETs jump to line rate</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-9-0-get-bandwidth.svg" alt="GET bandwidth by value size, comparing Valkey 8.1.8 in orange with Valkey 9.0.4 in green"></p>
<p>Every size from 1 MB to 64 MB reads at 190-201 Gbps. The 8 MB p99 falls from 112 to 24 ms, and a 64 MB read drops from just under a second to 230 ms. Writes do not move because the ingest path was never copy-bound. That ceiling waits for 9.1.</p>
<h3 id="valkey-91-more-headroom-for-writes">Valkey 9.1: more headroom for writes</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-9-1-set-bandwidth.svg" alt="SET bandwidth by value size, comparing Valkey 9.0.4 in dark green with Valkey 9.1.0 in light green"></p>
<p>With reads pinned at the network limit, the gain surfaces on writes. The 8 MB write rises from 134 to 166 Gbps, a 24% gain, while values from 12 MB to 64 MB gain 11-19%.</p>
<h2 id="the-short-run-held">The short run held</h2>
<p>To make sure the 15-second runs were not catching a lucky window, I reran every 8 MB GET and SET cell for 15 minutes. Valkey 9.1 is a good example: GET held 200.8 Gbps and SET held 165.5 Gbps, right on top of the original 201 and 166 Gbps results. Nothing sagged as the runs went on.</p>
<picture>
  <source media="(max-width: 640px)" srcset="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/long-run-stability-mobile.svg">
  <img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/long-run-stability.svg" alt="Valkey 9.1 GET and SET throughput holding steady over 15 minutes" loading="lazy">
</picture>
<h2 id="detailed-results">Detailed results</h2>
<h3 id="reads-get">Reads (GET)</h3>
<p>GET-only, 32 connections, 100% hit rate.</p>
<h4 id="get-bandwidth-gbps">GET bandwidth (Gbps)</h4>



























































<table><thead><tr><th>Value size</th><th align="right">7.2.13</th><th align="right">8.1.8</th><th align="right">9.0.4</th></tr></thead><tbody><tr><td>1 MB</td><td align="right">32</td><td align="right">183</td><td align="right">200</td></tr><tr><td>2 MB</td><td align="right">35</td><td align="right">55</td><td align="right">193</td></tr><tr><td>4 MB</td><td align="right">26</td><td align="right">45</td><td align="right">197</td></tr><tr><td>8 MB</td><td align="right">21</td><td align="right">33</td><td align="right">200</td></tr><tr><td>12 MB</td><td align="right">21</td><td align="right">32</td><td align="right">191</td></tr><tr><td>16 MB</td><td align="right">21</td><td align="right">32</td><td align="right">201</td></tr><tr><td>32 MB</td><td align="right">21</td><td align="right">32</td><td align="right">196</td></tr><tr><td>64 MB</td><td align="right">22</td><td align="right">31</td><td align="right">191</td></tr></tbody></table>
<p>Note: Valkey 9.1 reads also hold line rate, so the table stops at 9.0.</p>
<h4 id="get-p99-latency-ms">GET p99 latency (ms)</h4>



























































<table><thead><tr><th>Value size</th><th align="right">7.2.13</th><th align="right">8.1.8</th><th align="right">9.0.4</th></tr></thead><tbody><tr><td>1 MB</td><td align="right">13</td><td align="right">2.2</td><td align="right">2.4</td></tr><tr><td>2 MB</td><td align="right">20</td><td align="right">19</td><td align="right">5.1</td></tr><tr><td>4 MB</td><td align="right">58</td><td align="right">41</td><td align="right">9.4</td></tr><tr><td>8 MB</td><td align="right">174</td><td align="right">112</td><td align="right">24</td></tr><tr><td>12 MB</td><td align="right">244</td><td align="right">166</td><td align="right">56</td></tr><tr><td>16 MB</td><td align="right">329</td><td align="right">238</td><td align="right">56</td></tr><tr><td>32 MB</td><td align="right">531</td><td align="right">489</td><td align="right">116</td></tr><tr><td>64 MB</td><td align="right">948</td><td align="right">1,020</td><td align="right">230</td></tr></tbody></table>
<h3 id="writes-set">Writes (SET)</h3>
<p>SET-only, 32 connections.</p>
<h4 id="set-bandwidth-gbps">SET bandwidth (Gbps)</h4>




































































<table><thead><tr><th>Value size</th><th align="right">7.2.13</th><th align="right">8.1.8</th><th align="right">9.0.4</th><th align="right">9.1.0</th></tr></thead><tbody><tr><td>1 MB</td><td align="right">51</td><td align="right">201</td><td align="right">201</td><td align="right">201</td></tr><tr><td>2 MB</td><td align="right">53</td><td align="right">200</td><td align="right">200</td><td align="right">201</td></tr><tr><td>4 MB</td><td align="right">55</td><td align="right">200</td><td align="right">200</td><td align="right">201</td></tr><tr><td>8 MB</td><td align="right">23</td><td align="right">137</td><td align="right">134</td><td align="right">166</td></tr><tr><td>12 MB</td><td align="right">23</td><td align="right">135</td><td align="right">139</td><td align="right">166</td></tr><tr><td>16 MB</td><td align="right">24</td><td align="right">137</td><td align="right">139</td><td align="right">165</td></tr><tr><td>32 MB</td><td align="right">24</td><td align="right">139</td><td align="right">134</td><td align="right">159</td></tr><tr><td>64 MB</td><td align="right">24</td><td align="right">130</td><td align="right">134</td><td align="right">149</td></tr></tbody></table>
<h4 id="set-p99-latency-ms">SET p99 latency (ms)</h4>




































































<table><thead><tr><th>Value size</th><th align="right">7.2.13</th><th align="right">8.1.8</th><th align="right">9.0.4</th><th align="right">9.1.0</th></tr></thead><tbody><tr><td>1 MB</td><td align="right">8.7</td><td align="right">3.3</td><td align="right">3.2</td><td align="right">3.2</td></tr><tr><td>2 MB</td><td align="right">17</td><td align="right">6.5</td><td align="right">6.7</td><td align="right">5.3</td></tr><tr><td>4 MB</td><td align="right">35</td><td align="right">12</td><td align="right">13</td><td align="right">14</td></tr><tr><td>8 MB</td><td align="right">164</td><td align="right">41</td><td align="right">49</td><td align="right">34</td></tr><tr><td>12 MB</td><td align="right">289</td><td align="right">62</td><td align="right">57</td><td align="right">52</td></tr><tr><td>16 MB</td><td align="right">357</td><td align="right">77</td><td align="right">70</td><td align="right">60</td></tr><tr><td>32 MB</td><td align="right">1,580</td><td align="right">118</td><td align="right">143</td><td align="right">107</td></tr><tr><td>64 MB</td><td align="right">2,580</td><td align="right">237</td><td align="right">213</td><td align="right">206</td></tr></tbody></table>
<h2 id="the-setup">The setup</h2>
<p><strong>Machines.</strong> Two <code>c8gn.16xlarge</code> instances with Graviton4 and 200 Gbps networking in the same availability zone and cluster placement group, running Amazon Linux 2023.</p>
<p><strong>Server.</strong> We tested the official <code>valkey/valkey</code> Docker image at versions <code>7.2.13</code>, <code>8.1.8</code>, <code>9.0.4</code>, and <code>9.1.0</code>. We also measured <code>8.0.9</code>. It matched <code>8.1.8</code> within run-to-run noise at every size, so the tables show 8.1 as the 8.x column. Valkey 7.2 ran with the same flags as the newer versions. A control with <code>io-threads-do-reads yes</code> changed only the 1 MB GET cell, from 32 to 41 Gbps. Each version ran in a fresh container with host networking:</p>
<pre><code class="language-sh">docker run --network host --cpuset-cpus 8-23 \
  --ulimit nofile=32768:65536 valkey/valkey:&#x3C;version> \
  --save '' --appendonly no --io-threads 16 \
  --protected-mode no --maxmemory 60gb
</code></pre>
<p>Persistence was off. The 16 threads in Valkey’s <code>io-threads</code> count were one main thread plus 15 I/O workers. The process was pinned to cores 8-23 so it never fought the kernel for the cores doing network interrupt work.</p>
<p><strong>Interrupts.</strong> <code>irqbalance</code> was off on both machines. The ENA NIC was configured with four combined queues and their IRQs pinned to cores 0-3. Without this step, results wander from run to run as the kernel shuffles interrupts onto whatever cores the server or client threads happen to be using. If you benchmark at these speeds, pin your IRQs first and thank yourself later.</p>
<pre><code class="language-sh"># Run on both machines. ens50 is the ENA interface name on these instances.
sudo systemctl stop irqbalance
sudo ethtool -L ens50 combined 4
i=0
for irq in $(grep ens50 /proc/interrupts | awk '{print $1}' | tr -d ':'); do
  echo $((i%4)) | sudo tee /proc/irq/$irq/smp_affinity_list > /dev/null
  i=$((i+1))
done
</code></pre>
<p><strong>Client.</strong> We used <a href="https://github.com/cachecannon/cachecannon">valkey-lab</a>, built on cachecannon, with 16 worker threads pinned to cores 4-19, 32 connections, and pipeline depth 1. Reads and writes were measured in separate passes. Read passes prefilled the keyspace and ran at a 100% hit rate. Each short-run cell was a 15-second measurement after a five-second warmup, over a 500-key keyspace with 16-byte keys.</p>
<p>We kept concurrency at 32 connections because AWS caps a single TCP flow at roughly 9.5 Gbps. We verified 9.53 Gbps with iperf3. Saturating a 200 Gbps NIC requires spreading the load across flows.</p>
<p>One invocation per cell, with <code>-s</code> swept across the value sizes:</p>
<pre><code class="language-sh"># GET pass: prefill the 500-key keyspace, then measure at a 100% hit rate.
valkey-lab -h &#x3C;server> -c 32 -P 1 -t 16 --cpu-list 4-19 \
  -n 500 -r 100:0 --prefill --warmup 5s -d 15s \
  -s &#x3C;value_bytes> --key-size 16

# SET pass.
valkey-lab -h &#x3C;server> -c 32 -P 1 -t 16 --cpu-list 4-19 \
  -n 500 -r 0:100 --warmup 5s -d 15s \
  -s &#x3C;value_bytes> --key-size 16
</code></pre>
<p><strong>Scope of the result.</strong> This is a single-node, no-TLS, persistence-disabled test with 32 closed-loop connections, pipeline depth 1, and a 100% hit rate. Most table cells are one 15-second measurement after warmup. The results describe this rig and workload. Broader production claims need repeated runs and additional configurations.</p>
<p><em>We benchmarked whole releases rather than individual changes, so we treat the pull requests named in this post as leading explanations rather than proof of causality.</em></p>
<h2 id="what-this-means-if-you-run-valkey">What this means if you run Valkey</h2>
<p>Valkey 9.x delivers roughly 6x the 64 MB read bandwidth of 8.x while cutting p99 from just under a second to about 230 ms on this single-node, no-TLS test. Our earlier mixed-workload test also found far less collateral latency for small requests when a large read arrived.</p>
<p>The practical takeaway is narrower than “Valkey is always faster.” Valkey 8.1 is essentially flat on this workload, Valkey 9.0 changes the large-GET path, and Valkey 9.1 improves large SETs on this rig. If large values matter to your workload, test 9.x with your object-size distribution, concurrency, TLS, and persistence settings rather than extrapolating from a small-object benchmark, or from this one.</p>
<p>Valkey has consistently improved performance, memory efficiency, and availability at scale. Each version brings about a new set of improvements driven by issues faced by real users in production. A vibrant community where nobody is incentivized to withhold features for the sake of revenue and everyone is incentivized to chase continuous improvement is what makes Valkey truly special. Effort is additive. Consistency is multiplicative.</p>
<p>Serving megabyte-sized objects at wire speed is what I’ve been spending a lot of time on lately. If you are wrangling larger objects, let’s talk. <a href="https://valkey.io/slack/">Join me on Valkey Slack</a>.</p>]]></content:encoded>
  </item>
  <item>
    <title>The concurrency cliff is a memory limit</title>
    <link>https://www.gomomento.com/jp/blog/the-concurrency-cliff-is-a-memory-limit/</link>
  <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
    <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
  <category><![CDATA[Inference]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/the-concurrency-cliff-is-a-memory-limit/</guid>
    <description><![CDATA[<p>Latency looks fine on p50 until the KV cache fills. Then p99 jumps 15x in one step while your average dashboard shows nothing wrong. The cliff is memory.</p>]]></description>
    <content:encoded><![CDATA[<p>Add concurrent users to an inference server and latency usually creeps up. KV cache serving does not creep. It holds flat, then falls off a cliff.</p>
<p>On a single L4 running a Qwen3-4B coding agent, the server holds 12 concurrent sessions at p99 under 2.6 seconds. Add two more and p99 jumps to 39 seconds, a 15× increase in a single step. The cliff is the moment the KV cache fills.</p>
<h2 id="an-agentic-coding-workload-on-commodity-hardware">An agentic coding workload on commodity hardware</h2>
<p>Our setup is a single <code>g6.4xlarge</code> EC2 instance with one NVIDIA L4 GPU (24 GB VRAM), running <a href="https://huggingface.co/Qwen/Qwen3-4B">Qwen3-4B</a> with FP8 weights on vLLM 0.20.2, automatic prefix caching (APC) enabled.</p>
<p>The workload models a lightweight coding agent mid-task. Each session opens with a 10,000-token shared system prompt (repository context plus agent instructions), followed by a 12-turn conversation where each turn appends about 1,000 tokens of unique context (tool call inputs, code snippets, responses). At turn 12, total context reaches roughly 22,800 tokens per session. Output is capped at 75 tokens per turn. The workload is heavily input-dominated, as agentic workloads tend to be.</p>
<pre><code>System prompt:          10,000 tokens  (shared across sessions → APC cached)
Per-session turns:      ~12,800 tokens across 12 turns  (unique per session)
Output per turn:        75 tokens
—————————————————
Total at turn 12:       ~22,800 tokens
</code></pre>
<p>APC caches the system prompt once and amortizes its KV cost across all sessions (the blocks still occupy GPU memory, but only one copy exists). Within a session, APC also caches the growing turn history. Turn n+1 extends the exact prefix from turn n, so each turn only prefills the new ~1,000 tokens, as long as the prior turns’ KV blocks survive in cache. But the per-session turn history is unique (different code, different tool outputs) and cannot be shared across sessions. When concurrency pressure forces eviction of a session’s blocks, the next request in that session has to re-prefill the full accumulated history (up to ~12,800 tokens at turn 12). That re-prefill cost drives the TTFT cliff.</p>
<p>We swept concurrency from 1 to 48 sessions, measuring TTFT at each level, across three KV cache precisions, fp16, fp8, and TurboQuant 4-bit. We also ran best-case (all context cached) and worst-case (all context re-prefilled) bounds to bracket where realistic performance should land.</p>
<h2 id="the-concurrency-cliff">The concurrency cliff</h2>
<p>The chart below shows TTFT percentiles (p50, p95, p99) and throughput across the full concurrency sweep with fp8 KV cache. The Y axis is logarithmic, and even on a log scale the cliff is sharp.</p>
<p><img src="https://www.gomomento.com/blog/2026-06-26_automatic-prefix-caching/ttft-percentiles.avif" alt="chart: TTFT p50/p95/p99 + throughput vs concurrent sessions, fp8 KV">
<em>TTFT p50, p95, and p99, with throughput on the right axis, versus concurrent sessions. fp8 KV cache, Qwen3-4B FP8 on a g6.4xlarge (L4). The shaded band marks the 12 to 16 session collapse, and the dashed line marks the KV cache filling at 14c.</em></p>
<p>Below 12 concurrent sessions, p99 TTFT stays under 2.6 seconds and throughput climbs to a peak of 1.36 req/s. At 14 sessions, p99 jumps to 38.9 seconds in a single step, a 15× increase, and throughput drops 23 percent at the same time.</p>
<p>At 12 sessions, everything fits in KV cache. At 14, filling it forces eviction of another session’s blocks. That evicted session has to re-prefill its full 12,800 tokens of unique context at its next turn, and the cascade collapses latency.</p>
<p>The cliff sits between 12 and 14 sessions, and throughput never recovers past the peak at 12c. For a 2-second p99 SLA, the safe operating point is 6 concurrent sessions.</p>
<h3 id="after-the-cliff-p50-lies-to-you">After the cliff, p50 lies to you</h3>
<p>Above the cliff, p50 and p99 live in different regimes. At 20c, p50 is 2.2 seconds, which looks manageable, while p99 is 44.8 seconds, which is not. The distribution is bimodal, because APC creates two populations of requests. Lucky requests hit warm cache entries for their session context, prefill only the latest turn, and finish fast. Unlucky requests arrive after their session’s blocks were evicted, re-prefill the full 12,800 tokens, and take 30 to 50 seconds under load.</p>
<p>The p50 reflects the lucky cohort and the p99 reflects the unlucky one. A single TTFT average is meaningless past the cliff. You have to look at the tail to see the failure.</p>
<p>A few points from the sweep show the whole shape, flat through 12 sessions then the cliff at 14 and the widening p50/p99 gap past it.</p>





























































<table><thead><tr><th>Sessions</th><th>TTFT p50</th><th>TTFT p95</th><th>TTFT p99</th><th>req/s</th></tr></thead><tbody><tr><td>1</td><td>388</td><td>487</td><td>493</td><td>0.42</td></tr><tr><td>6</td><td>1.06s</td><td>1.43s</td><td>1.75s</td><td>1.18</td></tr><tr><td>10</td><td>1.19s</td><td>1.89s</td><td>2.31s</td><td>1.33</td></tr><tr><td>12</td><td>1.25s</td><td>2.16s</td><td>2.60s</td><td>1.36</td></tr><tr><td>14</td><td>1.29s</td><td>8.93s</td><td>38.9s</td><td>1.05</td></tr><tr><td>20</td><td>2.25s</td><td>25.5s</td><td>44.8s</td><td>0.68</td></tr><tr><td>48</td><td>5.74s</td><td>40.5s</td><td>43.2s</td><td>0.89</td></tr></tbody></table>
<h2 id="kv-cache-precision-moves-the-knee">KV cache precision moves the knee</h2>
<p>Running the same workload with 16-bit KV cache (vLLM defaults to the model’s dtype, bf16 for Qwen3, when <code>--kv-cache-dtype</code> is not set) halves the token capacity, and the knee shifts left in proportion.</p>

































<table><thead><tr><th>KV dtype</th><th>Bits/element</th><th>Token capacity</th><th>Knee (sessions)</th><th>2s p99 ceiling</th></tr></thead><tbody><tr><td>bf16/fp16</td><td>16</td><td>~89K</td><td>~8</td><td>~4</td></tr><tr><td>fp8</td><td>8</td><td>~178K</td><td>~14</td><td>~6</td></tr><tr><td>TurboQuant 4-bit</td><td>~4.2</td><td>~275K</td><td>~23 (est.)</td><td>pending</td></tr></tbody></table>
<p>The fp16 to fp8 shift is confirmed, with fp16 knees at about 8 and fp8 at about 14, a 1.75× shift for a 2× capacity increase. The slight compression below 2× is expected, since KV management overhead and block table fragmentation consume some of the headroom regardless of precision.</p>
<p><img src="https://www.gomomento.com/blog/2026-06-26_automatic-prefix-caching/ttft-p99.avif" alt="chart: TTFT p99, fp8 KV vs fp16 KV, knees marked">
<em>TTFT p99 for fp8 KV against fp16 KV, same workload and GPU. The markers sit at the observed knees, 8 sessions for fp16 and 14 for fp8. The tq4 knee near 23 is a capacity-based estimate, pending the experiment.</em></p>
<p>Quantization buys concurrency headroom directly. Halving KV precision from fp16 to fp8 nearly doubles how many concurrent sessions fit before the cliff. TurboQuant 4-bit (about 3.8× fewer bytes per element than fp16, partially offset by the lower <code>gpu_memory_utilization</code> it needs for autotuning scratch space) predicts a knee at about 23c, roughly 3× more concurrent sessions than fp16.</p>
<p>The accuracy tradeoff may be small. KV cache quantization at 4-bit typically reports low single-digit perplexity impact, though the exact effect depends on model and task. The knee shifts from about 8 sessions (fp16) to about 14 (fp8) to an estimated 23 (tq4), roughly 3× more sessions before eviction onset, from the same GPU.</p>
<h2 id="best-case-worst-case-realistic">Best case, worst case, realistic</h2>
<p>To separate the latency budget that is fundamental (prefill compute) from the part that is avoidable (cache misses), we ran two controlled bounds alongside the realistic workload.</p>
<p>In the best case (<code>miss_rate=0.0</code>), every request hits the same cached content. APC holds the full 12,800-token session context, so only about 200 unique tokens need prefilling, which is perfect KV utilization.</p>
<p>The worst case (<code>miss_rate=1.0</code>) gives every request a unique prefix that breaks APC for the user context. The 10K system prompt still hits the cache, but all ~12,800 tokens of per-session turn history are re-prefilled on every request. Every miss lands at peak session depth (turn 12), forcing the maximum re-prefill cost each time.</p>
<p><img src="https://www.gomomento.com/blog/2026-06-26_automatic-prefix-caching/ttft-p50.avif" alt="chart: TTFT p50, best vs realistic vs worst, log scale">
<em>TTFT p50 for best (fully cached), realistic (natural APC), and worst (always re-prefilled), on a log scale. The band between best and worst is the envelope any real workload lands in.</em></p>
<h3 id="what-the-bounds-tell-us">What the bounds tell us</h3>
<p>At a single concurrent session, with zero contention, the three bounds separate cleanly.</p>

























<table><thead><tr><th>Workload</th><th>TTFT p50</th><th>What’s happening</th></tr></thead><tbody><tr><td>Best (cached)</td><td>57 ms</td><td>Only ~200 unique tokens prefilled; rest is cached</td></tr><tr><td>Realistic (APC)</td><td>388 ms</td><td>System prompt cached; 12,800 unique tokens prefilled</td></tr><tr><td>Worst (evicted)</td><td>2,500 ms</td><td>System prompt cached; ~12,800 user tokens re-prefilled at peak depth every request</td></tr></tbody></table>
<p>On an absolute scale, realistic (388 ms) is much closer to best (57 ms) than to worst (2,500 ms). But realistic is still 7× slower than best. That gap is the cost of prefilling about 12,800 tokens of per-session unique context on each request. APC removes the system prompt cost, but the per-session turn history still has to be computed.</p>
<p>The gap between realistic and worst is about miss depth. In the realistic workload, cache misses happen at any turn. A session evicted at turn 3 re-prefills about 3,000 tokens, while eviction at turn 12 costs about 12,800. The worst case forces every miss to peak session depth, paying the maximum re-prefill on every request. Real traffic produces a distribution of miss depths, which is why realistic latency stays close to best.</p>
<p>The best-case result is the surprising one. With perfectly cached session context, the L4 handles more than 48 concurrent sessions within a 2-second p99 SLA. The 6 session realistic ceiling is the cost of per-session context uniqueness, the turn histories that cannot be shared. It is not a GPU compute limit.</p>
<p>The worst case grows linearly at about 2.1 seconds per additional concurrent session, reaching 102 seconds at 48c. Throughput saturates at 0.32 req/s from 6 sessions onward. The GPU is fully consumed re-prefilling 12,800 tokens per request, and extra concurrency just lengthens the queue.</p>
<h2 id="why-the-knee-is-where-it-is">Why the knee is where it is</h2>
<p>The L4 has 24 GB of VRAM, but far less than that is available for KV cache. The memory that actually holds KV cache is roughly half of raw VRAM.</p>
<h3 id="where-the-memory-goes">Where the memory goes</h3>
<p>vLLM’s <code>gpu_memory_utilization</code> was set to 0.9 for the fp8 and fp16 experiments, reserving about 21.6 GB. After model weights, CUDA graph capture, activation tensors, and block table overhead, about 13 GB remains for KV cache. The TurboQuant experiment used 0.8 (it needs about 2 GB of extra scratch for torch.inductor autotuning at startup), leaving about 10.6 GB.</p>
<h3 id="kv-cache-per-token">KV cache per token</h3>
<p>Qwen3-4B uses GQA with 36 layers, 8 KV heads, and head_dim 128. The per-token KV cache size depends on precision.</p>
<pre><code>2 (K+V) × 36 layers × 8 KV heads × 128 head_dim × bytes_per_element

FP16/BF16: ... × 2 bytes = 147,456 bytes/token  → ~89K tokens in ~13 GB
FP8:       ... × 1 byte  =  73,728 bytes/token  → ~178K tokens in ~13 GB
TQ4:       ~0.53 B effective (4-bit + quantization metadata)
           =  ~38,700 bytes/token  → ~275K tokens in ~10.6 GB
</code></pre>
<h3 id="the-capacity-arithmetic-with-apc">The capacity arithmetic (with APC)</h3>
<p>With APC, the 10,000-token system prompt is stored once and shared. Only the per-session unique context (about 12,800 tokens at peak depth) needs its own blocks.</p>
<pre><code>FP8 KV
Token budget:     ~178K
Shared prefix:     10K (1×)
Available:        ~168K
Per-session:      ~12.8K
Max sessions:   168K / 12.8K ≈ 13      Observed knee: ~14c

FP16 KV
Token budget:      ~89K
Shared prefix:     10K (1×)
Available:         ~79K
Per-session:      ~12.8K
Max sessions:    79K / 12.8K ≈ 6       Observed knee: ~8c

TQ4 KV (estimated)
Token budget:     ~275K (0.8 util)
Shared prefix:     10K (1×)
Available:        ~265K
Per-session:      ~12.8K
Max sessions:   265K / 12.8K ≈ 21      Predicted knee: ~23c
</code></pre>
<p>The arithmetic predicts the knees within 1 to 2 sessions of the observed values. The slight overshoot (observed 14 sessions against predicted 13) is because sessions are not all at peak depth at once. Earlier turns have smaller contexts, which buys a few extra sessions before capacity runs out.</p>
<p>In this setup, the concurrency cliff is a memory limit. The binding constraint is how many sessions’ KV caches fit in VRAM at once. The best-case bound supports this. With perfect caching, the same GPU handles more than 48 sessions within 2-second p99. Compute, scheduling, and continuous batching also contribute, but memory capacity sets the ceiling.</p>
<h2 id="how-to-find-the-knee-for-your-workload">How to find the knee for your workload</h2>
<p>The knee location depends on three variables. Available KV cache memory is total VRAM minus model weights, CUDA graphs, activations, and fragmentation, typically about half of raw VRAM, and vLLM reports the exact number at startup. Per-session unique context is the total session tokens at peak depth, minus any shared prefix cached by APC. KV precision is the bytes per element, where halving it from fp16 to fp8 to 4-bit roughly doubles token capacity at each step and shifts the knee right.</p>
<p>The estimate is max concurrent sessions ≈ (token capacity − shared prefix) / per-session unique context.</p>
<p>For this setup (Qwen3-4B, L4, 22.8K-token agentic sessions with a 10K shared prefix), the arithmetic predicts about 13 sessions (fp8) and 6 (fp16). The observed knees are about 14 and 8. The arithmetic gives a first-order estimate, and a concurrency sweep gives the precise number. The gap between estimate and observation comes from session depth staggering, block fragmentation, and APC reuse patterns.</p>
<p>Different workloads shift each variable. A single-turn QA workload with 2K tokens per session has a much higher knee. A code review agent with 50K-token inputs has a much lower one. A GPU with more VRAM (A100, H100) raises the budget. The method is the same. Estimate the budget, divide by per-session cost, then verify with a sweep.</p>
<h2 id="what-this-means-for-deployment">What this means for deployment</h2>
<p>Know your KV budget before you set your concurrency limit. Below the knee you get the best throughput with stable latency and effective caching. Above it you get worse throughput, worse latency, and a wasted APC investment.</p>
<p>KV quantization is a direct concurrency multiplier. On this L4, switching from fp16 to fp8 KV cache moves the 2-second p99 SLA ceiling from about 4 to 6 (50 percent more sessions) and the eviction knee from about 8 to 14 (75 percent more sessions). The gain is a direct consequence of halving the bytes per KV element. Quantization buys memory, and memory buys concurrency.</p>
<p>Monitor tail latency, not averages. After the cliff, p50 looks manageable while p99 is catastrophic. The bimodal distribution means some users get sub-second responses while others wait 40-plus seconds, and an average-based dashboard hides it until users complain.</p>
<p>If some of your sessions are latency-tolerant background work, run them off the interactive path. They do not need to compete for cache memory, and keeping them off it frees KV budget for the sessions that need low TTFT.</p>
<p>Several caveats temper the numbers. These experiments use synthetic token content, not real code. The workload has a fixed 12-turn structure, while real agent sessions vary widely in depth. Poisson arrivals do not capture bursty agentic traffic, where agents send follow-up requests immediately. p99 at high concurrency is noisy, since with about 200 requests per run it is only the second-worst request. Chunked prefill, not enabled here, could smooth the knee transition. The numbers are specific to a single L4 with Qwen3-4B. Larger models, multi-GPU setups, and different context lengths shift the absolute numbers while the pattern holds.</p>
<p>KV cache behaves like a systems problem, and the concurrency knee is where that meets a specific GPU, a specific model, and a specific workload shape. The math is simple. The discipline is running it before production tells you the hard way.</p>]]></content:encoded>
  </item>
  <item>
    <title>Your KV cache benchmark is “hi hi hi”</title>
    <link>https://www.gomomento.com/jp/blog/your-kv-cache-benchmark-is-hi/</link>
  <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
    <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
  <category><![CDATA[Inference]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/your-kv-cache-benchmark-is-hi/</guid>
    <description><![CDATA[<p>Default KV cache benchmarks run on repetitive "hi hi hi" text, which makes compression and transfer look far better than they do on real workloads.</p>]]></description>
    <content:encoded><![CDATA[<p>Before you commit to a KV cache offloading system, you benchmark it to make sure it performs well below your SLA. You see that it has excellent compression and cheap transfers. Seems like an easy win.</p>
<p>But there’s some trouble with what standard KV cache benchmarks run on.</p>
<p>LMCache ships with a <a href="https://docs.lmcache.ai/getting_started/benchmarking.html">long-document benchmark</a> for measuring KV cache offloading performance. Run it without a corpus file and it generates documents like this:</p>
<pre><code class="language-python">warmup_prompts = [
  str(i) + " " + " ".join(["hi"] * args.document_length)
  for i in range(args.num_documents)
]
</code></pre>
<p>A 10,000-token document comes out looking a little underwhelming.</p>
<p><code>0 hi hi hi hi hi hi hi hi hi hi hi hi hi hi hi …</code></p>
<p>Technically speaking, it is the 10K token count you were looking for, but it doesn’t represent a real 10K token workload.</p>
<p>KV cache systems do not run on token count alone. Compression ratios, activation patterns, transfer sizes, and cache behavior all depend on the shape of the input. Two documents of the same length can be two entirely different workloads.</p>
<p>Unfortunately, much of the current KV cache ecosystem is benchmarked on synthetic inputs that look nothing like the workloads people run.</p>
<h2 id="the-benchmark-is-not-representative">The benchmark is not representative</h2>
<p>The default benchmark document contains a numeric identifier and the token “<em>hi</em>” repeated thousands of times.</p>
<p>Transformers do not produce identical activations for repeated tokens. Positional encoding, attention mixing, and residual connections keep every position distinct. To an LLM, <em>distinct</em> and <em>varied</em> are not the same thing. Repeating a single token produces far more regular activation patterns than diverse text does.</p>
<p>Compression improves, transfer sizes shrink, and cache behavior becomes easier to predict with non-varied workloads. The benchmark is measuring <em>something</em>, but it is not measuring a realistic production workload.</p>
<p>Benchmarking KV cache offloading with “hi hi hi” is like benchmarking a database with <code>SELECT 1</code>. The numbers come back fast, but they do not tell you much about real workloads.</p>
<h2 id="the-difference-shows-up-immediately">The difference shows up immediately</h2>
<p>We compared the default benchmark document against a realistic medical document using <a href="https://huggingface.co/Qwen/Qwen3-8B-FP8">Qwen3-8B-FP8</a>. Both ran about 10,000 tokens. The synthetic one carried two unique tokens, a token diversity of 0.02 percent. The medical one carried 1,329, or 13.3 percent. The token count is the same. Everything else is different.</p>
<p>Token diversity affects activation patterns. Activation patterns affect tensor value distributions. Tensor distributions affect compression ratios and transfer sizes. A benchmark built from highly repetitive inputs can diverge sharply from one built on realistic text.</p>
<h2 id="compression-behaves-differently-too">Compression behaves differently too</h2>
<p>We first ran into this while building a Valkey-backed KV cache connector. We followed the LMCache tutorials and used dummy weights, and compression ratios looked excellent. Then we switched to Qwen3-8B-FP8 with real trained weights, and the results were night and day.</p>
<p>Real model weights produce tensor values that behave more like high-entropy floating-point data. General-purpose compression still helps, but the gains are smaller than they appear with dummy weights. Repetitive inputs create more structured activation patterns that compress more effectively, making the benchmark results appear better than they actually are.</p>
<h2 id="so-we-built-a-more-realistic-corpus">So we built a more realistic corpus</h2>
<p>To see how KV cache systems behave under representative inputs, we built a corpus of 30 long-form documents across medical and legal domains. We wanted documents that resemble the structure, formatting, vocabulary, and variability that real systems process.</p>
<p><img src="https://www.gomomento.com/blog/2026-06-24_kv-cache-benchmark/corpus.avif" alt="Example medical and legal document text"></p>
<p>Medical documents averaged 14.2 percent token diversity. Legal documents averaged 9.4 percent. Both are hundreds of times more diverse than the synthetic baseline at 0.02 percent.</p>
<p>The corpus holds 300,000 tokens and was generated with Qwen3-8B-FP8. The corpus and generation scripts are <a href="https://github.com/momentohq/kvcache-corpus">open source</a> for those interested.</p>
<h2 id="benchmarking-the-workload-you-actually-have">Benchmarking the workload you actually have</h2>
<p>While token count is easy to generate, it can also be the least informative. Match the diversity, structure, and vocabulary of the text your system serves, and the compression ratios and transfer sizes start to represent values you can trust.</p>
<p>The input comes first. Before ranking cache connectors, compression schemes, or storage backends, run them on inputs that look like your traffic. Compare them on “hi hi hi” and you are ranking them on a workload nobody runs.</p>
<p>Before you evaluate your next KV cache offloading system, be sure to ask “Are the benchmark documents representative of the workload I actually run?”</p>]]></content:encoded>
  </item>
  <item>
    <title>vLLM's Hash Chain and Why Prefix Caching Is Still Prefix Caching</title>
    <link>https://www.gomomento.com/jp/blog/prefix-caching-is-still-prefix-caching/</link>
  <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
    <pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate>
  <category><![CDATA[Inference]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/prefix-caching-is-still-prefix-caching/</guid>
    <description><![CDATA[<p>Automatic prefix caching, hash chains, radix trees. The data structures get cleverer, but we haven’t moved past shared prefixes.</p>]]></description>
    <content:encoded><![CDATA[<p>Automatic Prefix Caching sounds like it should solve a bigger problem than plain prefix caching. Requests are hashed, cache entries are discovered automatically, and shared work never has to be tracked by hand. It looks like a system that can find reusable computation wherever it appears.</p>
<p>But it’s still the same type of reuse we’ve always had. Shared prefixes are reusable. Shared content that is not a prefix is not. That rule sets the biggest limit on what today’s inference infrastructure can reuse.</p>
<p>Most of the recent work in KV caching has focused on finding prefixes more efficiently. Hash chains, radix trees, and automatic discovery all improve the mechanics of reuse. But they don’t change what can be reused. The workloads where it succeeds, and where it falls short, show why.</p>
<h2 id="when-prefix-caching-is-enough">When prefix caching is enough</h2>
<p>Reusing the KV cache only pays off when requests share computation, so the question is how much of a real workload today’s infrastructure can reuse.</p>
<p>For many agentic workloads, the answer is often “enough”. Stable system prompts, tool definitions, and conversation scaffolding create long shared prefixes, and inference engines are good at finding and reusing them.</p>
<p>Plenty of workloads share more than just prefixes, though. RAG pipelines are where it breaks down. The system prompt stays fixed while the retrieved documents change from request to request. Two requests might carry the same five documents in a different order. The meaning is almost identical. The token sequence is not, and the cache matches on the token sequence. Same content, different positions, no reuse.</p>
<h2 id="why-prefix-caching-remains-prefix-bound">Why prefix caching remains prefix-bound</h2>
<p>vLLM’s <a href="https://docs.vllm.ai/en/v0.20.1/features/automatic_prefix_caching/">Automatic Prefix Caching</a> uses content hashing to remove the need for explicit prefix tracking. The KV cache is divided into fixed-size blocks, and each block gets a SHA-256 hash. The hash for block N folds in the hash of every preceding block along with the content of the current block. Chained together, every block fingerprints not only its own content but the entire token history before it.</p>
<p>When a request arrives, vLLM hashes each block-sized chunk of input and checks whether a matching block already exists. Matching blocks reuse previously computed KV cache. Missing blocks trigger fresh computation. Lookup is effectively constant-time, eviction is straightforward, and fixed-size blocks map cleanly onto GPU memory.</p>
<p>But because each block hash depends on every block before it, one divergence changes every hash that follows. Take two requests that share the first 3,000 tokens and split at token 3,001. The block holding token 3,001 hashes differently, and so does every block after it. Reuse stops at the point of divergence. The system discovers shared prefixes on its own, and it can’t discover shared content that appears once requests have diverged.</p>
<p>Reuse happens only at block boundaries. If two requests share 1,000 tokens and a block holds 16, vLLM reuses 62 whole blocks, or 992 tokens, and recomputes the remaining 8. For long prefixes that waste is negligible. For short or irregular shared segments it’s more substantial.</p>
<p>There is no matching inside a block, either. Two blocks that share 15 of their 16 tokens still hash to completely different values, so reuse is all-or-nothing at the block level. These are reasonable tradeoffs that keep the implementation simple and fast, but they still leave you with the same limitation: reuse follows exact prefix structure.</p>
<p>A different cache structure might help. <a href="https://docs.sglang.io/">SGLang</a> takes that route. Instead of hashing fixed-size blocks, it keeps cached state in a <a href="https://www.lmsys.org/blog/2024-01-17-sglang/">radix tree indexed by token sequences</a>. When a request arrives, the runtime walks the tree and finds the longest matching cached prefix, and matches can fall on arbitrary token boundaries rather than fixed block ones. That helps workloads with variable-length turns, branching conversations, and irregular prefix lengths.</p>
<p>The radix tree still searches for the longest shared prefix, though. Once two requests diverge, the content they share later in the sequence stays out of reach. SGLang improves how prefixes are discovered, but it does not extend reuse past prefixes.</p>
<h2 id="beyond-prefixes">Beyond prefixes</h2>
<p>Prefix caching tells us that KV reuse clearly works. The more open question now is how much reuse survives divergence. Today’s runtimes are tuned for shared prefixes. The next generation goes after shared segments, cache repair, and the reuse that prefix matching cannot reach.</p>
<p>So much of the current research now focuses on cache repair and segment-level reuse. The goal has shifted from proving that KV reuse is valuable to recovering the work that today’s prefix-based systems still leave behind.</p>]]></content:encoded>
  </item>
  <item>
    <title>Disaggregation makes KV cache a system primitive</title>
    <link>https://www.gomomento.com/jp/blog/disaggregation-makes-kv-cache-a-system-primitive/</link>
  <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
    <pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate>
  <category><![CDATA[Inference]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/disaggregation-makes-kv-cache-a-system-primitive/</guid>
    <description><![CDATA[<p>Prefill and decode want different hardware. Separate them and the KV cache becomes the state that connects the two, which turns it from an implementation detail into a system design concern.</p>]]></description>
    <content:encoded><![CDATA[<p>Inference is scaling faster than the serving architectures around it. Prefill and decode are often treated as one pipeline, but they are fundamentally different workloads.</p>
<p>Prefill is compute-heavy, with an order of magnitude more arithmetic intensity than decode. Decode is sensitive to memory bandwidth and to latency. Prefill wants high-FLOPS accelerators. Decode wants large, fast memory. Put both on the same GPU and you tune it for one profile while it absorbs the other, and under load prefill interferes with decode. Neither phase gets the hardware it would choose.</p>
<h2 id="the-cache-becomes-the-connection">The cache becomes the connection</h2>
<p>Disaggregation separates these two phases. Prefill nodes do prefill, decode nodes do decode, each on hardware suited to its bottleneck.</p>
<p>Separation removes the interference, but it opens a gap. Inside one accelerator, prefill flows straight into decode and the intermediate state never leaves the chip. Pull the two onto different machines and that state, the KV cache, has to be handed across. Prefill produces it, decode consumes it, and disaggregation turns it into the object that connects them.</p>
<p>On a single node, the KV cache is largely an implementation detail the inference engine manages for you. But once prefill and decode are separate systems, the cache is the connection between them, and every request depends on getting it from one to the other.</p>
<h2 id="what-the-split-asks-of-the-cache">What the split asks of the cache</h2>
<p>Once the cache has to travel between machines, it takes on the requirements of any object moving through a distributed system.</p>
<p>The cache has to move from prefill nodes to decode nodes, which makes transfer latency, serialization format, and network bandwidth first-order concerns. It has to land in the right place, so deciding which decode node receives which cache turns routing into a scheduling problem. It has to expire, which raises the question of who evicts an entry and when, work the engine handled in a colocated system and that now needs coordination. And it has to live somewhere across GPU memory, host memory, NVMe, and remote storage, each tier with its own latency and capacity tradeoffs.</p>
<p>These are distributed systems problems that present themselves whenever you separate storage and compute. What was colocated becomes independently addressable, connected by a transfer layer. The techniques for solving them, placement, routing, eviction, and tiered storage, are well understood. What is new is applying them to the KV cache inside inference serving.</p>
<h2 id="from-implementation-detail-to-system-design">From implementation detail to system design</h2>
<p>The major inference stacks are already built around this split. <a href="https://developer.nvidia.com/dynamo">NVIDIA Dynamo</a> describes disaggregated inference as a prefill engine that computes the prefill phase and generates KV cache, hands that cache to a decode engine, and lets the decode engine run the decode phase. AWS is building the split into its infrastructure, with <a href="https://aws.amazon.com/machine-learning/trainium/">Trainium</a> for compute-heavy prefill, and the <a href="https://www.cerebras.ai/blog/cerebras-is-coming-to-aws">Cerebras partnership</a> likely points the same way, since wafer-scale SRAM suits memory-bound decode.</p>
<p>In each of these, the KV cache is the object the tiers hand between them. On a single node it was an optimization you could run, skip, or tune, and the architecture did not care. Disaggregation moves those same problems up a level, from implementation details the engine used to hide to system design concerns the architecture has to own. The cache now has to be transferred, routed, stored, and expired across machines, and the whole system is built around getting it from prefill to decode.</p>
<p>What began as transient state inside a single request becomes the handoff between two systems. Once we have that handoff, cache management becomes a first-class primitive of the architecture.</p>]]></content:encoded>
  </item>
  <item>
    <title>KV Caching Pays Off Under Load</title>
    <link>https://www.gomomento.com/jp/blog/kv-caching-pays-off-under-load/</link>
  <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
    <pubDate>Tue, 16 Jun 2026 07:00:00 GMT</pubDate>
  <category><![CDATA[Inference]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/kv-caching-pays-off-under-load/</guid>
    <description><![CDATA[<p>KV cache starts out as an implementation detail of inference. As inference systems evolve, it is becoming a first-class systems primitive.</p>]]></description>
    <content:encoded><![CDATA[<p>KV caching looks like a bad trade on paper. Memory, complexity, and operational surface area, all spent to shave a few percent off a request.</p>
<p>The benchmarks do not rescue it.</p>
<p>We’ve seen teams leave it at that. KV cache is necessary inside a single forward pass, so you keep it for the life of the request and move on. Keeping it alive beyond that, reused across requests, starts to sound like a serving-layer luxury. You picture the memory it would pin, the eviction logic, the extra moving parts in a stack that is already hard enough to operate. Set that against a few percent of latency and the trade does not look worth making.</p>
<p>Understandably so. Run the numbers on a single request and long-lived KV caching underwhelms. We ran them, and at first it was a very unflattering story. But a single request is the wrong unit to judge this on, and once you measure at production load the economics turn. </p>
<h2 id="the-single-request-savings-are-bounded">The single-request savings are bounded</h2>
<p>In one of our <a href="https://ollama.com/library/qwen3:30b-a3b">Qwen3-30B-A3B</a> runs, a 1K input / 512 output request came in around 135 ms TTFT and about 2.5 seconds of total request latency. TTFT carries scheduling and queueing overhead on top of raw prefill compute, so call the prefill portion roughly 100 ms. Erase it completely, the most a perfect cache hit can do, and you save about 4 percent of the request. Treat that as the ceiling, not the everyday case.</p>
<p>As context grows, so does prefill’s share of the total request latency. At 16K input / 512 output, TTFT was about 769 ms of 3,200 ms total, which puts it near 24 percent. That is a real step up from the 1K case. The input/output ratio affects prefill more than the context length does. KV cache earns the most when a request carries a large input and returns a small output, because prefill is then a bigger share of the bill. In scenarios where you have short input and long output, decode takes over while the cache has little room to help.</p>
<p>On its own, a 4 or 24 percent share looks modest. Latency is measured at a target throughput, GPU capacity is scarce, and as throughput climbs, utilization, queueing, and pipeline stalls increase the cost of redoing prefill. So in production, prefill becomes a capacity and tail-latency problem once thousands of requests compete for the same accelerators.</p>
<p>So the skepticism is fair, for the single request. If the only question is whether one cache hit meaningfully cuts one request’s latency, the answer is often no, and it depends on the input/output ratio and how much of the request prefill actually owns. At that level, KV caching is not an automatic win.</p>
<p><img src="https://www.gomomento.com/blog/2026-06-16_kv-caching-pays-off-under-load/prefill-vs-decode.avif" alt="Prefill vs decode ratios"></p>
<p>Large-input, small-output workloads are getting more common, not less. Agentic workloads are multi-turn by nature. Context grows as chat history, tool-call results, and retrieval chunks pile up, so each new turn carries a larger input against a small output. Exactly the type of workload where prefill dominates and where reusing the cache has the most to give.</p>
<p>For application teams, reusing the cache shows up as lower TTFT, tighter p95, and lower cost per request. For the platform teams running the GPUs, it shows up as higher utilization and more capacity per dollar. </p>
<h2 id="expensive-to-hold-hard-to-reuse">Expensive to hold, hard to reuse</h2>
<p>That said, the KV cache is not free to keep. Hold it in GPU memory and it competes with active inference for the scarcest space you have. Move it to host memory and it is cheaper but still bounded, with DRAM prices trending the wrong way. Push it out to remote memory or storage and you take on transfer latency, placement problems, and more operational surface. KV caching is a bet. You are spending scarce memory on the wager that future requests reuse the work you are holding.</p>
<p>But that’s only half the problem. Even when you are willing to pay for the memory, the reuse you get back is limited. The production-friendly option today is prefix caching: if a later request begins with the exact same prefix, the engine reuses the KV cache already computed for it. The rule is strict, exact prefix match or nothing. Plenty of real workloads share meaning without sharing a prefix. Reordered retrieval chunks, varying tool results, and shifting user context carry the same semantic content in different positions, and none of it counts as a hit, so hit rates suffer. </p>
<p>It might seem like KV caching has a lot going against it. The single-node latency win is bounded. The memory cost is high. The reuse model is narrow. Evaluate it as an isolated optimization on one node and the honest question is whether the complexity earns its place. In isolation, often it does not. But isolation is the wrong frame because inference is not the system it was when those objections were formed. Each of them was measured against a single node running prefill and decode together, holding a cache that was large and expensive to keep. Two things have shifted since. The first is structural, in where prefill and decode run. The second is economic, in what the cache costs to hold and to move. Each one undercuts a different piece of the case against. </p>
<h2 id="inference-is-becoming-a-distributed-systems-problem">Inference is becoming a distributed systems problem</h2>
<p>Prefill and decode are not the same kind of work. Prefill is compute-heavy. Decode is sensitive to memory bandwidth and to latency. Put both on the same accelerator and you force a compromise on one to serve the other. Split them, and you create a clean boundary between two workloads that want different things. The KV cache, however, has to cross that boundary.</p>
<p>When prefill and decode live on different hardware, the KV cache becomes a first-class distributed systems primitive, something you transfer, place, and manage a lifecycle for. NVIDIA Dynamo and the disaggregated stacks coming out of AWS and Cerebras are building the split into the infrastructure itself, which is what forces developers to think about how the KV cache moves, where it lives, and how long it stays alive.</p>
<h2 id="the-economics-are-starting-to-move">The economics are starting to move</h2>
<p>The second shift is economic. The KV cache itself is getting more efficient to store and to move. A surprising amount of recent model progress is really KV cache innovation, and the last 18 months have been striking. </p>
<p><a href="https://arxiv.org/abs/2405.04434">DeepSeek-V2 and V3</a> introduced <a href="https://mccormickml.com/2025/04/26/inner-workings-of-mla/">Multi-head Latent Attention (MLA)</a>. MLA compresses keys and values into a shared low-rank latent vector before anything gets cached. For V3, that takes the per-token cache from roughly 16,384 scalar dimensions under standard multi-head attention down to 576, a 512-dimensional latent plus 64 dimensions for decoupled RoPE. Against an MHA baseline that is about a 28x reduction. Against the GQA baseline most modern models already use, it is smaller, roughly 4 to 8x depending on group size, and MLA gets there while holding MHA-level quality, which GQA gives up.</p>
<p><a href="https://qwen.ai/blog?id=qwen3.5">Qwen 3.5</a> goes a different way with a <a href="https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/08_deltanet/README.md">Gated DeltaNet hybrid</a>. It replaces 75 percent of its attention layers with Gated DeltaNet linear attention, layers that hold a fixed-size state matrix, 128 by 128 per head, and update it incrementally with each token. The state does not grow with sequence length. Only the remaining 25 percent, full softmax attention with GQA, still needs a traditional KV cache. At long contexts, where the KV cache usually dominates memory, this removes most of the growth. The payoff scales with context: substantial at 256K tokens, modest at 1K, where a fixed-size state costs about what a small KV cache would anyway.</p>
<p><a href="https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/">TurboQuant and PolarQuant</a>, from Google at ICLR 2026, take yet another angle. Instead of changing the attention mechanism, they quantize the KV cache itself to 3 or 4 bits per coordinate with no measurable accuracy loss on standard benchmarks. PolarQuant rotates vectors with a random orthogonal matrix so the coordinates follow a known distribution, then applies an optimal Lloyd-Max scalar quantizer, and QJL adds a 1-bit residual correction. At 4 bits the paper reports up to 8x faster attention on an H100. At 3 bits, roughly 6x memory reduction.</p>
<p>The exact numbers depend on baselines and configurations, but the direction is obvious. MLA shrinks the cache dramatically. Hybrid architectures such as Qwen’s Gated DeltaNet reduce cache growth across much of the network. Quantization approaches such as TurboQuant reduce memory requirements further without changing the model architecture. Different tradeoffs, same trend: the object is getting smaller. </p>
<p>Memory cost is the usual objection to KV caching, but this recent work almost makes it moot. Shrink the cache by 6x to an order of magnitude and the economics look very different. More entries fit in the same budget. Transfers from remote memory, SSD, or another node get faster. There’s a misconception that the cache has to become trivially small. But it only has to get small enough that the economics cross over for the workloads people run in production.</p>
<p>The storage hierarchy is changing as well. The previous thought was if the KV cache is not in GPU memory, it is too slow to matter. That is getting harder to say. Fast interconnects and local NVMe continue to improve. Now the question is whether moving or loading the cache can free the decode GPU from repeated prefill work and keep it pointed at the latency-sensitive part of the pipeline. If a storage-backed cache lowers prefill pressure and keeps accelerator capacity on decode, it can pay off quickly.</p>
<p>The workload is the main success driver. For a given model and context length, the system weighs the time to recompute prefill against the time to move the cache over the network, the time to read it from SSD, the cost of reserving the memory or storage, and the odds the cache gets reused at all. When transfer or load time comes in well under recompute time and reuse is likely enough, the cache earns its place. When it does not, the cache is a cost with no return.</p>
<h2 id="prefix-caching-under-load">Prefix caching under load</h2>
<p>We saw this behavior in an experiment we recently ran. We used Qwen3-1.7B on an L40S, a 10K-token prompt, and the number of shared prefix tokens varied from 0 to 10K across several concurrency levels. As the shared prefix grows, the vLLM prefix cache hit ratio climbs from 0 to 1. Throughput and request latency were monitored at each concurrency level.</p>
<p><img src="https://www.gomomento.com/blog/2026-06-16_kv-caching-pays-off-under-load/plot_cache_hit_ratio.avif" alt="Plot graph of cache hit ratio"></p>
<p><em>Prefix cache hit ratio grows linearly with shared prefix length, from 0 to 10K tokens against the fixed 10K-token prompt.</em></p>
<p><img src="https://www.gomomento.com/blog/2026-06-16_kv-caching-pays-off-under-load/plot_rps.avif" alt="Plot graph of shared prefix and concurrency"></p>
<p><em>Throughput rises as the shared prefix grows and is sharpest at high concurrency, where skipping repeated prefill lets the system sustain more requests per second.</em></p>
<p><img src="https://www.gomomento.com/blog/2026-06-16_kv-caching-pays-off-under-load/plot_request_latency.avif" alt="Plot graph of request latency"></p>
<p><em>Mean request latency falls as reuse increases, with the high-concurrency settings improving most.</em> </p>
<p>At low concurrency, a higher hit ratio helps, but only a little. At high concurrency, the same increase results in a much larger system effect. Requests per second climb sharply as more of the prompt comes from cache, and mean latency decreases along with it, dropping fastest at the higher concurrency levels. That is the behavior one would expect if KV caching is a systems optimization rather than a single-request latency trick.</p>
<p>The experiment is a small model on a single node, so it does not prove the disaggregated-architecture argument on its own. But it does verify that prefix cache hits remove prefill work from the serving path, and the system-level benefit grows with concurrency. The disaggregation thesis is that this gets stronger when prefill and decode run on separate hardware and the KV cache moves between them as a first-class object.</p>
<h2 id="prefix-caching-already-works-for-the-right-workloads">Prefix caching already works for the right workloads</h2>
<p>Prefix caching generally has a narrow sharing mode. For agentic workflows it fits more naturally than you might expect, though the reason it fits changes by category of context.</p>
<p>System prompts are the easy case. They are stable across requests, they sit at the front of the prompt, and are a textbook prefix hit. An agent making a series of tool calls against the same backend reuses the same 2K to 8K token system prompt on every request. A multi-turn conversation with a fixed system prompt reuses the whole instruction block. A code-generation agent with stable repository context reuses the project description and file summaries. For this category, cross-request caching is straightforward.</p>
<p>Other kinds of context ask for more care. Chat history grows and shifts from turn to turn. Tool-call exemplars get reordered or swapped. Retrieval chunks change with every RAG query. These often share material across requests <em>without sharing an exact prefix</em>, so the effectiveness of the cache comes down to how much of the context is positionally stable (which prefix caching handles), versus variable (which needs something like CacheBlend to unlock). </p>
<h2 id="research-for-a-better-solution">Research for a better solution</h2>
<p>For messier patterns, like RAG with retrieval chunks that vary by query or tool results that differ between calls, two requests can share a great deal of material without sharing the exact same prefix, and classic prefix caching returns a miss in those cases even when most of the computation could have been reused.</p>
<p><a href="https://arxiv.org/abs/2405.16444">CacheBlend</a> is one of the research directions in this area. It is exploring the idea of <em>cache repair</em>, which takes a semantically similar cached entry to what the current request needs, and selectively recomputes only the parts that differ. If repair is cheap enough, individual caches become reusable across more requests and hit rates rise without spending more memory.</p>
<p>This is still in open research, it’s not solved yet. No major inference framework ships chunk-level KV reuse today. The selective recomputation carries its own latency, quality preservation depends on the workload, and the methodology for measuring these tradeoffs is still maturing. But the direction is promising. More flexible cached entries raise the effective hit rate inside the same memory budget.</p>
<p>Prefix caching answers “<em>does caching work</em>?” for a growing number of workloads. The open question is how much of the rest can be brought into the cacheable regime, and cache repair is where we are working that out.</p>
<h2 id="from-per-request-state-to-systems-primitive">From per-request state to systems primitive</h2>
<p>If you evaluate KV caching as an isolated optimization on a single node, it doesn’t make much sense. The memory cost is high and prefix-based reuse is limited.</p>
<p>But the architecture underneath is changing. Disaggregated prefill and decode create the right interface. Better attention mechanisms shrink the object you have to store and move. Faster networks and SSDs reduce transfer costs. Cache repair could push reuse beyond strict prefixes. Scarce GPU capacity makes repeated prefill work harder to justify.</p>
<p>Together, these shifts are turning the KV cache from a temporary intermediate state into an inference systems primitive.</p>]]></content:encoded>
  </item>
  <item>
    <title>Beyond the Goals, Three Ways Momento Scales the Football World Cup in Real Time</title>
    <link>https://www.gomomento.com/jp/blog/beyond-the-goals-three-ways-momento-scales-the-football-world-cup-in-real-time/</link>
  <dc:creator><![CDATA[Lionel Bringuier]]></dc:creator>
    <pubDate>Wed, 03 Jun 2026 16:14:30 GMT</pubDate>
  <category><![CDATA[Media & Entertainment]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/beyond-the-goals-three-ways-momento-scales-the-football-world-cup-in-real-time/</guid>
    <description><![CDATA[<p>When the world&amp;#039;s biggest sporting event kicks off, every millisecond matters. Learn how Momento helps broadcasters and sports platforms deliver faster, smarter, and more secure fan experiences at global scale.</p>]]></description>
    <content:encoded><![CDATA[<p><img src="https://www.gomomento.com/wp-content/uploads/2026/06/Lionel-Banner-1024x512.png" alt=""></p>
<p><em><a href="https://www.gomomento.com/wp-content/uploads/2026/06/FIFA-infographic.pdf">Don’t have five minutes? This infographic summarizes the key takeaways from this blog.</a></em></p>
<p> The FIFA World Cup 2026 is not just a tournament for Momento. It is a live fire test of what real-time data at global scale actually means.</p>
<p>In stadiums and on sofas, hundreds of millions of fans will see goals, cards, and heartbreak. Behind the scenes, three of Momento’s largest customers will be doing something just as intense: pushing a real-time data platform to the limit across live origination, content protection, and AI-powered personalization.</p>
<p>Beneath these vastly different workloads lies the same fundamental principle that decisions must be made instantly, at global scale, with zero excuses.</p>
<h2 id="what-do-we-mean-by-momento-is-a-real-time-data-platform">What Do We Mean by “Momento is a Real-Time Data Platform”?</h2>
<p>“Data platform” is one of those phrases that can mean anything from a gigantic static data warehouse to a firehose of events in flight. When we say Momento is a real-time data platform, we mean something specific as we combine:</p>
<ul>
<li><strong>A sub-millisecond in-memory data engine</strong>: This is the RAM-cache based data plane that serves reads and writes in less than a millisecond, even under massive load.</li>
<li><strong>An intelligent control plane</strong>: This layer automatically handles sharding, scaling, partitioning, and hot-key management, so app teams do not have to be distributed systems experts</li>
</ul>
<p>In practice, customers can treat Momento like a simple API to store and retrieve state in real time: segments and manifests, concurrency counters, user events, AI embeddings, and more. They describe their data model and policy. Momento makes it fast, durable, and observable.</p>
<p>The FIFA World Cup is a perfect way to show what that actually looks like when the stakes are highest. Think of it as a hat-trick of real-time use cases.</p>
<h2 id="1-a-live-origin-that-just-does-not-flinch">1/ A Live Origin That Just Does Not Flinch</h2>
<p><strong>Who:</strong> A large UK broadcaster holding FIFA rights
<strong>Problem:</strong> Their live origin service (AWS Elemental MediaStore) was deprecated before the World Cup. They needed the same low latency, failover behavior and observability, at World Cup scale, without rewriting their encoder, packager, video player or CDN configurations.</p>
<p><strong>How Momento helps:</strong><a href="https://www.gomomento.com/solutions/momento-media-storage/">Momento Media Storage</a> is their new live origin. It gives them:</p>
<ul>
<li><strong>Predictable, low latency:</strong> Consistent latency for reads and writes at the live edge.</li>
<li><strong>Granular TTL control:</strong> TTL settings on manifests/segments to preserve automatic failover capabilities.</li>
<li><strong>Per-Service Limits and Metrics:</strong> Observability and metrics are built in so high traffic events don’t impact the rest of their 24/7 channels.</li>
</ul>
<p>For viewers, nothing “looks” different. For their ops teams, the origin is now actively developed, supported, faster, and ready for tens of millions of concurrent fans.</p>
<p>🔗 Deep dive <a href="https://www.gomomento.com/blog/a-new-live-streaming-origin-built-for-global-scale/">on How a major UK broadcaster moved its World Cup live origin to Momento ↗</a></p>
<h2 id="2-content-protection-through-concurrency-tracking">2/ Content Protection Through Concurrency Tracking</h2>
<p><strong>Who:</strong> A major US broadcaster holding FIFA rights
<strong>Problem:</strong> Pirates leverage stolen accounts and run illegal restreaming operations on top of the broadcaster’s own CDN. Traditional DRM and short-lived tokens verify a device can decrypt, but are completely blind to whether an account is behaving like a bot farm. Content rights holders are concerned about piracy and are asking broadcasters to implement server-side control over account behavior to shut down the illegal streams at the source.</p>
<p>**How Momento helps:**The broadcaster runs a server-side <a href="https://www.gomomento.com/wp-content/uploads/2025/04/Momento_Concurrency_Overview_OnePager_FINAL_3.17.25.pdf">concurrency service</a> backed by Momento. For each stream, a lightweight verification loop executes three steps:</p>
<ul>
<li><strong>Receive the heartbeat:</strong> The player calls a Momento-powered service with identifiers for the subscriber’s account, their device, the content they want to access, their geography, etc.</li>
<li><strong>Enforce stream limits:</strong> Verify limits on concurrent streams per account and per event.</li>
<li><strong>Deliver instant decisions:</strong> Make allow/deny decisions in single-digit milliseconds, inline with playback.</li>
</ul>
<p>If one account suddenly powers hundreds of devices on the same game, the system can automatically shut it down, protecting revenue, CDN bills, and QoE for legitimate fans.</p>
<p>🔗 Details on the architecture and anti-leeching approach:
<a href="https://www.gomomento.com/blog/stop-cdn-leeching-with-concurrency-tracking/">Stop CDN Leeching with Concurrency Tracking ↗</a></p>
<h2 id="3-ai-powered-personalized-feeds-for-a-sports-app">3/ AI-Powered Personalized Feeds for a Sports App</h2>
<p><strong>Who:</strong> A US-based popular Sports content Super App
<strong>Problem:</strong> Just ahead of the World Cup, this content provider wanted to relaunch their app, with a brand new User Experience. Beyond the traditional “click to watch” from their editorial content, they needed a real-time, personalized feed that mixes editorial pieces, highlights, YouTube clips, and content from popular social networks, tuned to each fan’s behavior as it happens. Think of it as a TikTok “For You Page”, for Sports.</p>
<p>**How Momento helps:**Momento serves as the real-time event collection and embedding layer, making real-time machine learning models visible to the app users at massive scale:</p>
<ul>
<li><strong>Dynamic content ingestion:</strong> New content from newsrooms, creators, social networks, and athletes is ingested and turned into <a href="https://youtu.be/DlKiWkHtsXE?si=LJAcq01lzxwOCJlb">AI embeddings</a>.</li>
<li><strong>Real-time signal streaming:</strong> User signals including emoji reactions, comments, watch time, and scroll depth stream into Momento in real time.</li>
<li><strong>Sub-millisecond personalization:</strong> Recommendation services query Momento’s sub-millisecond data plane to match fresh content to each fan’s evolving interests.</li>
</ul>
<p>The result: A feed that feels instantly relevant and keeps improving as fans interact with their personalized feed, at the scale of millions of concurrent users.</p>
<p>🔗 Video explainer on real-time embeddings with Momento:
<a href="https://youtu.be/DlKiWkHtsXE?si=LJAcq01lzxwOCJlb">Momento AI embeddings &#x26; event collection explainer ↗</a></p>
<h2 id="lets-watch-football-not-infrastructure">Let’s Watch Football, Not Infrastructure</h2>
<p>Three very different workloads, all powered by the same underlying platform: a low latency data plane with an intelligent control layer that simplifies real-time data usage for application teams.</p>
<p>When the first match kicks off, the fans will be watching football, not infrastructure. But inside control rooms, NOCs, and product teams, our customers will know. They will see cleaner dashboards, stronger protections, and faster feedback loops. They will see a real-time data platform doing exactly what it was built to do.</p>
<p>We are excited to be part of their World Cup story. And once the final whistle blows, these capabilities will not disappear. They will become the new baseline for what fans expect from live sports, everywhere.</p>]]></content:encoded>
  </item>
  <item>
    <title>A New Live Streaming Origin Built for Global Scale</title>
    <link>https://www.gomomento.com/jp/blog/a-new-live-streaming-origin-built-for-global-scale/</link>
  <dc:creator><![CDATA[Lionel Bringuier]]></dc:creator>
    <pubDate>Thu, 28 May 2026 20:14:36 GMT</pubDate>
  <category><![CDATA[Media & Entertainment]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/a-new-live-streaming-origin-built-for-global-scale/</guid>
    <description><![CDATA[<p>A major UK broadcaster rebuilt its live streaming origin on Momento ahead of the FIFA World Cup 2026.</p>]]></description>
    <content:encoded><![CDATA[<p><img src="https://www.gomomento.com/wp-content/uploads/2026/05/Untitled-design-8-1024x512.png" alt=""></p>
<p>It’s the world’s most watched sport. In the US, it’s soccer; everywhere else, it’s football; and in the UK, it’s nearly a religion. In the run-up to the FIFA World Cup 2026, a large UK-based broadcaster made a big bet: they migrated their live streaming origin off from the deprecated AWS Elemental MediaStore service to Momento, and they did it in time to serve tens of millions of fans.</p>
<p>This wasn’t a simple lift and shift. The broadcaster’s live stack is a mature, battle-tested system that has evolved over years of 24×7 live linear channels, major tournaments and peak news moments. Replacing the core media storage and origin layer under that stack meant Momento had to meet an exacting bar on latency, reliability and operational visibility, while continuing to support an existing fleet of encoders and packagers, CDNs, and control-plane tooling.</p>
<p>This post walks through the existing architecture, adapting Momento as a drop-in replacement for AWS Elemental MediaStore, and then hardening the infrastructure for the World Cup and beyond.</p>
<h2 id="the-starting-point-live-origin-at-scale">The Starting Point: Live Origin at Scale</h2>
<p>The broadcaster runs a large portfolio of 24×7 simulcast channels plus frequent pop-up live events in both HD and UHD, delivered via DASH and HLS. Streams are published into multiple AWS regions for redundancy, and each stream fans out across multiple CDNs.</p>
<p>To give a rough estimate of the data volume, their live origin manages 40+ live 24×7 channels, with some additional seasonal pop-up channels (up to 40 concurrent ones for large sport events). The live channels are encoded in H.264 and HEVC, with typically 8 to 11 video encoding profiles and 4 audio tracks in the ABR ladder. Segments and manifests are pushed into a media storage service deployed in two AWS regions, with cross-AZ replication in each region. On the playback side, CDNs pull from that origin, either directly or via an internal cache concentration layer to distribute traffic across providers and geographies.</p>
<p>At first, the broadcaster evaluated moving the origin to a standard S3 bucket. However, their architecture and operations depended on certain capabilities that would not be guaranteed.</p>
<p>When AWS Elemental MediaStore became deprecated, the broadcaster faced a classic “you have to rebuild the airplane while flying it” challenge: replace a key service in their workflow without rewriting their packagers, CDN configs, video players, or control planes, and without introducing new failure modes at the worst possible time, a year ahead the global football tournament. What were the tenets for the new service?</p>
<ul>
<li><strong>Fast live-edge reads and writes</strong> from London-based clients: tight time-to-first-byte and time-to-last-byte for both PUT (publication from the encoders) and GET (playback).</li>
<li><strong>Transient data policies</strong> to aggressively expire stale manifests, forcing automatic failover to backup origins when the primary stopped publishing.</li>
<li><strong>Per-container request limits</strong> to prevent one high-traffic service from drowning others.</li>
<li><strong>Per-container access policies and CORS</strong> for secure origin access from CDNs and packagers.</li>
<li><strong>Detailed access logging and metrics</strong> for publication latency, error codes, empty object publications, and regional breakdowns.</li>
<li><strong>Lifecycle management</strong> to trim historical content and control storage costs for the content in the hot cache.</li>
<li><strong>Keep the durability of an S3-backed storage</strong> under the hood, but without compromising access latency consistency.</li>
</ul>
<h2 id="design-goal-a-drop-in-mediastore-replacement">Design Goal: A Drop-In MediaStore Replacement</h2>
<p>The joint design goal was straightforward to state but hard to achieve: “just swap the origin to <a href="https://www.gomomento.com/solutions/momento-media-storage/">Momento Media Storage</a> with minimal application changes, while preserving, or improving, the operational semantics we rely on today”.</p>
<p>Concretely, that translated into a few core requirements for Momento:</p>
<ul>
<li><strong>Equivalent HTTP surface area</strong>Keep the same style of HTTP PUT/GET semantics, 404/50x behavior for missing segments, and origin-side access control via headers and tokens.</li>
<li><strong>Performance parity or better from London</strong>For UK-based clients, Momento had to deliver GET and PUT latencies that matched or beat their existing origin across both eu-west-1 (Dublin) and eu-west-2 (London).</li>
<li><strong>Configurable object TTLs to emulate transient data policies</strong>Instead of path-based lifecycle rules on containers, the broadcaster wanted fine-grained TTL control per object class (e.g., manifests vs segments) to preserve their model where a stale manifest can trigger failover.</li>
<li><strong>Operational observability that matched their current dashboards</strong>Publication latency, error codes, throttling, per-path metrics, and near-real-time access logs had to remain available for the broadcaster’s existing monitoring and alerting workflows.</li>
<li><strong>Sensible multi-tenant safety rails</strong>Per-service request limits and regional SLAs needed to be enforced in a way that matched their mental model from the previous platform.</li>
</ul>
<h2 id="today-we-are-ready-for-kick-off">Today: We Are Ready for Kick-Off</h2>
<p>The journey from early performance tests to full production lasted almost a year. During that time, the broadcaster and Momento ran a substantial battery of tests:</p>
<ul>
<li><strong>Distributed publication and playback</strong> across dozens of channels in parallel.</li>
<li><strong>Comparative latency benchmarks</strong> from London-based clients to both eu-west-1 and eu-west-2, under varying bitrates and ladders.</li>
<li><strong>Load and failover drills</strong> to confirm that short manifest TTLs and 404 semantics still triggered the right automatic reactions in their CDN and player stack.</li>
<li><strong>SDK vs native HTTP tests</strong> to iron out any client-side inefficiencies and eliminate measurement artifacts.</li>
<li><strong>Reproducibility and automation,</strong> with a single orchestration layer for the whole video stack.</li>
<li><strong>Operational observability</strong> at every step of the workflow, that slots into their existing CloudWatch-based dashboards and alerting.</li>
</ul>
<p>Most importantly, they are now running on an origin layer that is actively developed, not deprecated, and that can evolve with their roadmap and future needs.</p>
<p>For now the focus is simple: when the referee blows the whistle to start the first World Cup match, tens of millions of fans across the UK and beyond will be watching through a new origin, built on Momento, and they won’t even notice the difference. And that’s exactly how it should be.</p>]]></content:encoded>
  </item>
  <item>
    <title>Introducing valkey-lab: Stop Guessing When Your Cache Hits Its Limit</title>
    <link>https://www.gomomento.com/jp/blog/introducing-valkey-lab-stop-guessing-when-your-cache-hits-its-limit/</link>
  <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
    <pubDate>Tue, 26 May 2026 20:14:51 GMT</pubDate>
  <category><![CDATA[Valkey]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/introducing-valkey-lab-stop-guessing-when-your-cache-hits-its-limit/</guid>
    <description><![CDATA[<p>Benchmarking a cache should be about finding how much load your system can sustain before latency SLOs break. Introducing valkey-lab.</p>]]></description>
    <content:encoded><![CDATA[<p><img src="https://www.gomomento.com/wp-content/uploads/2026/05/valkey-lab-1024x566.jpg" alt=""></p>
<p>Pop quiz: how many requests per second can your cache take before it stops meeting your latency SLO?</p>
<p>Chances are good you don’t know the answer, and that’s not a knock on you. It’s a genuinely hard number to come by. The standard tool for this is <a href="https://valkey.io/topics/benchmark/">valkey-benchmark</a>, and it’s great at exactly one thing: you point it at a server, you hammer it with some commands at full speed, and it prints a throughput number at the end. That number tells you the box is alive and roughly how fast it goes flat out.</p>
<p>From a production standpoint, that’s not as useful as it sounds.</p>
<p>How does the p999 hold up under an 80:20 read/write mix when half the requests land on hot keys? What does the tail look like at the rate you actually plan to run? How much headroom do you have before the SLO breaks? Was that latency spike at 10:32 a fluke or the ceiling? A single summary number printed after a sixty-second run can’t answer any of those, because it threw away the useful bits on the way to computing the average.</p>
<p><a href="https://github.com/cachecannon/cachecannon/blob/main/VALKEY-LAB.md">valkey-lab</a> was built to answer these questions.</p>
<p>It’s a high-performance Valkey and Redis benchmark that uses <code>io_uring</code> for kernel-bypassed I/O, per-connection pipelining, and multi-threaded workers. The defaults are deliberately familiar. Run it without any arguments:</p>
<p>📄</p>
<pre><code>valkey-lab

</code></pre>
<p>and you get a sixty-second run against <code>localhost:6379</code> with an 80:20 GET/SET ratio and a million keys. Same shape as the tool you already know, same short flags (<code>-h</code>, <code>-p</code>, <code>-c</code>, <code>-P</code>, <code>-r</code>). The interesting part starts once you begin asking harder questions.</p>
<h2 id="saturation-search-the-headroom-number-found-for-you"><strong>Saturation search: the headroom number, found for you</strong></h2>
<p>Here’s how you find your headroom number today, by hand. You run a benchmark at some rate, read the p999, decide it looks healthy, bump the rate, run it again, read it again. You do this five or ten times, squinting at each result, trying to find the rate where the tail skyrockets. Somewhere in that loop you lose track of which run had which config. Eventually you settle on a number you’re “pretty sure” is right and plan capacity around it. valkey-lab does that whole search for you with one command.</p>
<p>📄</p>
<pre><code>valkey-lab saturate --slo-p999 1ms -c 16 -P 32

</code></pre>
<p>A synthetic benchmark might say a cache can handle 2M requests per second. But when you add a realistic read/write mix, hot keys, and a warm cache, your p999 suddenly crosses your SLO at 1.2M. Technically speaking, the server is still processing 2M requests per second, but the usable ceiling is significantly lower.</p>
<p><code>saturate</code> starts issuing requests at whatever is provided in <code>--start-rate</code> (or 1000 if not provided) and multiplies the request rate by the <code>--step</code> factor on every step. The default step is <code>1.05</code>, so the load compounds over time. Each step holds its rate for a sample window, measures the full percentile spread, and checks it against your SLO. The moment a percentile crosses the line, the ramp stops and reports the last rate that held.</p>
<p>When a step fails, valkey-lab tells you <em>how</em> it failed, either throughput-limited or latency-exceeded, which helps you tune your clusters more accurately.</p>
<p>Throughput-limited means the server couldn’t generate the requested rate at all. It topped out below the target. That’s a capacity problem: you need more CPU, more shards, or a different topology.</p>
<p>Latency-exceeded means the server kept up with the rate, but the tail blew past the SLO. The server can sustain the requested rate, but something in the path is introducing tail spikes under load. Could be a hot key, a GC pause, a scheduler stall, network jitter. You fix that by chasing the spike, and adding hardware won’t help.</p>
<p>So based on your failure, your mitigation strategy varies wildly. And it would be impossible to know which one to pursue if all you had was the throughput number.</p>
<h2 id="averages-hide-the-interesting-failures"><strong>Averages hide the interesting failures</strong></h2>
<p>The next problem surfaces when the benchmark completes. Summary statistics hide the behavior you’re usually trying to find.  If your p999 was 312µs for fifty-nine seconds and 4.2 ms for one second, the run-level p999 still looks fine. The spike is the important part that you need to focus on.</p>
<p>valkey-lab streams one row per second with the full latency spread:</p>
<p><img src="https://www.gomomento.com/wp-content/uploads/2026/05/valkey-lab-1-1.jpg" alt=""></p>
<p>Every major percentile from p50 to p99.99 plus the max, the error count, and the cache hit rate, all per second. A spike that lasts one second appears as one row with a tall tail, a vast improvement over the executive summary at the end of a run. When you need it machine-readable instead, –output json gives you newline-delimited JSON you can pipe straight into something else, and –output quiet collapses the whole run to a single summary line. </p>
<h2 id="make-the-benchmark-look-like-your-workload"><strong>Make the benchmark look like your workload</strong></h2>
<p>There’s an important gotcha with the saturation number, or any benchmark number. A ceiling is only as good as the load that produced it, and the default load most tools run is unrealistic.</p>
<p>Think about what a stock benchmark actually does. It sends all reads, or close to it, because a 100% GET run posts the biggest number (or it’s the easiest to simulate). It picks keys uniformly at random, so every key is equally cold and nothing is ever hot. It runs flat out, measuring throughput at saturation. And it normally starts against an empty cache. Now think about your production traffic. It’s a read-write mix. It has hot keys, with a small fraction of the keyspace taking most of the requests. And the cache is warm. Every one of those differences takes away from the realism of the benchmark run.</p>
<p>valkey-lab addresses each one of these gaps. Set the real read-write split with -r so you’re measuring the write path your cache actually carries. Turn on –distribution zipf so a small fraction of keys receives most of the traffic, like production systems often do. Uniform access patterns avoid contention and hide the behavior of your actual hot paths. </p>
<p>Pin the load with –rate-limit to track latency at the rate you plan to run. And warm the cache with –prefill, or model a read-through cache that fills on miss with –backfill, so a GET benchmark measures hits the way production would.</p>
<p>📄</p>
<pre><code>valkey-lab --prefill -r 100:0 --distribution zipf -c 16 -P 32

</code></pre>
<p>Stack those and the ceiling you measure is a ceiling that meaningfully tracks production. There’s more depth when you need it, warmup tuning, RESP3, pinning workers to cores with –cpu-list, TLS, full TOML configs, but the move that matters is making the four big assumptions match your reality before you trust the number.</p>
<h2 id="getting-the-important-data-from-a-run"><strong>Getting the important data from a run</strong></h2>
<p>Now that we have realistic benchmark data, we have to make sure it’s useful after the run ends.</p>
<p><code>--parquet results.parquet</code> saves the full dataset to disk. It stores the full metric set per snapshot: the counters, the gauges, and the latency distributions as actual nanosecond histograms. Combine this with the visualization functionality in valkey-lab, and you have a rich experience that lets you dig into every tiny detail.</p>
<p>📄</p>
<pre><code>valkey-lab --parquet results.parquet
valkey-lab view results.parquet

</code></pre>
<p><code>view</code> opens an interactive dashboard against the file, with a synchronized time axis you can zoom and pan through dimensions like throughput, hit rate, error rate, and latency split out by GET, SET, and combined, all on a log scale. Scrub to the exact second p999 jumped and read every other metric in that same window. </p>
<p><img src="https://www.gomomento.com/wp-content/uploads/2026/05/valkey-lab-2-1024x508.jpg" alt=""></p>
<p>One use case for this is regression testing. Because every run is a Parquet file with the same schema, runs are directly comparable to each other. Benchmark before a Valkey upgrade and after, and the question “<em>did this move my tail latency</em>” is easily answered with a diff. The viewer is one way to read these files, but using your own queries is another easy way to act on changes in performance. <a href="https://duckdb.org/">DuckDB</a>, <a href="https://pandas.pydata.org/">pandas</a>, and <a href="https://pola.rs/">Polars</a> all read Parquet directly, so a few lines of SQL across a directory of runs is a regression suite for cache performance. Point DuckDB at a folder of recorded runs and let it compute peak throughput per file:</p>
<p>📄</p>
<pre><code>SELECT
  filename,
  max(responses_received) AS total_responses,
  max(request_errors)     AS errors
FROM read_parquet('runs/*.parquet', filename = true)
GROUP BY filename
ORDER BY filename;
</code></pre>
<p>That is a before-and-after table for every benchmark you have ever saved, built from data you already recorded. </p>
<p>Another use case for the Parquet output is root cause analysis. A spike on the latency chart tells you when something went wrong, not why. Point <code>view</code> at a <a href="https://github.com/iopsystems/rezolus">Rezolus</a> capture from the server or the client and it overlays system telemetry, CPU utilization, network, scheduler behavior, aligned to the same benchmark timeline. When a p999 spike lines up exactly with a scheduler stall or a network hiccup on the axis above it, you have your answer as simple as that.</p>
<h2 id="stop-guessing">Stop guessing</h2>
<p>Back to the pop quiz. The reason it’s so hard to answer is that traditionally the tool you use to measure max RPS reports a summary and throws the important bits away. valkey-lab changes the approach. It remembers the mix, the hot keys, the per-second tail, and records your runs so you can come back to them. The headroom number that used to take an afternoon of manual ramping is now a single command, and it comes with the failure mode attached so you know what to do about it.</p>
<p>valkey-lab is built on top of <a href="https://github.com/cachecannon/cachecannon">cachecannon</a>, inheriting its workload generation, saturation search, telemetry collection, and analysis capabilities. It needs Linux for io_uring (kernel 6.0+) and builds with Rust, under your choice of Apache-2.0 or MIT. Here is the whole getting-started path:</p>
<p>📄</p>
<pre><code>cargo install --path . --bin valkey-lab
valkey-lab saturate --slo-p999 1ms
</code></pre>
<p>Run that against a Valkey server and see what number comes back. Stop asking “<em>how fast can my cache go</em>” and start asking “<em>how fast can it go before my production workload breaks?</em>” That’s the number you capacity-plan around if you want predictable systems at 3 AM.</p>]]></content:encoded>
  </item>
  <item>
    <title>Why Snap Was Willing to Fork, and Why They Still Came Back</title>
    <link>https://www.gomomento.com/jp/blog/why-snap-was-willing-to-fork-and-why-they-still-came-back/</link>
  <dc:creator><![CDATA[Allen Helton]]></dc:creator>
    <pubDate>Thu, 21 May 2026 19:36:05 GMT</pubDate>
  <category><![CDATA[Caching]]></category>
  <category><![CDATA[Valkey]]></category>
    <guid isPermaLink="true">https://www.gomomento.com/jp/blog/why-snap-was-willing-to-fork-and-why-they-still-came-back/</guid>
    <description><![CDATA[<p>Snap ran 100% of their caching infrastructure on KeyDB, a Redis fork they acquired in 2022. Two years later, they moved to Valkey. Here&amp;#039;s why.</p>]]></description>
    <content:encoded><![CDATA[<p><img src="https://www.gomomento.com/wp-content/uploads/2026/05/fork-hero-1024x683.jpg" alt=""></p>
<p>I have no intention of ever forking a database. The amount of bravery and engineering mastery that goes into it scares me to no end. But Snap did. They committed to it so hard that they acquired the company building it, open sourced the entire commercial codebase, and ran 100% of their caching infrastructure on it for years. <a href="https://docs.keydb.dev/">KeyDB</a> powered Snapchat at a scale most companies can only dream of.</p>
<p>And then they migrated to Valkey anyway.</p>
<p>At <a href="https://www.unlockedconf.io/san-jose-replays">Unlocked San Jose</a>, Ovais Khan, Principal Software Engineer at Snap, walked through that migration. As interesting as it was to hear <em>how</em> they did it, it was all the more interesting to hear <em>why.</em> Why it happened, why it wasn’t worth staying on the fork, and why when they came back, they came back to Valkey. </p>
<h2 id="the-case-for-forking-in-2019"><strong>The case for forking in 2019</strong></h2>
<p>KeyDB started in 2019 as a project by John Sully and Ben Schermel at EQ Alpha Technology. The premise was simple. Redis ran a single-threaded event loop. Modern servers had 32, 64, 96 cores. To get peak throughput out of a single machine, you had to run a cluster of Redis nodes on it. That was wasteful, and Salvatore Sanfilippo, the creator of Redis, was on record <a href="https://antirez.com/news/126">arguing against changing it</a>: “<em>I/O threading is not going to happen in Redis AFAIK, because after much consideration I think it’s a lot of complexity without a good reason.</em>” Simplicity of the codebase was a value he was actively protecting. </p>
<p>KeyDB took the other side of that bet. It added real multithreading, with per-thread event loops and lock-based synchronization on shared state. It also added active-active replication and FLASH storage for cost-efficient large datasets. On the same hardware, it could move several times the operations per second that Redis could.</p>
<p>This is the textbook case for forking. The upstream project had made a deliberate architectural choice. That choice was the right one for them and the wrong one for a certain kind of user (Snap) who needed to push a single node harder. A fork was the only way forward.</p>
<p>By 2021, Snap was running KeyDB across enough of their caching infrastructure to want a permanent stake in it. They <a href="https://docs.keydb.dev/news/2022/05/12/keydb-joins-snap/">acquired the team in May 2022</a> and brought the formerly commercial KeyDB Pro features into the open source codebase under BSD-3. For about two years after that, all of Snap was running on KeyDB.</p>
<h2 id="what-forking-buys-you"><strong>What forking buys you</strong></h2>
<p>The benefits of forking are easy to articulate when you ship. Snap got features that were important for their specific operating model:</p>
<ul>
<li>Multithreaded command execution, which let them get more out of every node</li>
<li>Zone-aware read routing, which kept cross-AZ traffic down and cut data transfer costs considerably</li>
<li>Forkless background saves, which made snapshots predictable at high memory</li>
<li>Same-zone replica behavior that reduced timeout blast radius during upgrades</li>
</ul>
<p>These features weren’t going to make it into Redis on Snap’s timeline. The fork gave them room to build it as soon as they were ready.</p>
<p>As far as forking goes, that’s usually the part written in blog posts and talked about on the conference loop. You wanted a feature, the upstream said no, you built it yourself, and now it works. Forking feels like freedom.</p>
<h2 id="what-forking-costs-you"><strong>What forking costs you</strong></h2>
<p>Every change to upstream Redis after the fork point became a decision. Does it get ported over? Rewritten? Skipped? There’s a long tail at the end of whatever decision was made. Porting means you carry merge conflicts forever. Rewriting means you have two implementations of the same idea drifting apart. Skipping means your fork stops being a superset of upstream and starts being something else.</p>
<p>Ovais addressed this specifically in his talk. Snap could not easily move from KeyDB’s Redis 6.2 base to Redis 7.2. The cost of staying current with upstream had become high enough that they were stuck on a flavor of 6.2 while everyone else moved on. That meant they were also stuck without features the broader community had built on top of 7.2.</p>
<p>The same goes for the ecosystem. Every client library, operator, monitoring tool, and benchmark gets tested against upstream first. Your fork either matches upstream behavior closely enough that those tools just work, or it doesn’t, and you start maintaining your own.</p>
<p>While forking might have started off feeling like an accelerator, it quickly became a drag.</p>
<h2 id="the-redis-license-change"><strong>The Redis license change</strong></h2>
<p>In March 2024, Redis Ltd. changed the Redis license <a href="https://redis.io/blog/redis-adopts-dual-source-available-licensing/">from BSD-3 to a dual SSPL and RSALv2 model</a>. Neither license is OSI-approved. For any company offering Redis as a managed service, this was an immediate problem. AWS, Google Cloud, Oracle, and Ericsson responded by forking the last BSD release, Redis 7.2.4, and donating it to the Linux Foundation. Eight days after the license change, Valkey existed.</p>
<p>Up until then, the case for staying on KeyDB was obvious. The KeyDB team was inside Snap. The codebase was theirs. The performance was what they needed. </p>
<p>But Valkey made them pause. The project had open governance under the Linux Foundation, with a Technical Steering Committee across multiple companies and no single controlling vendor. It was BSD-licensed and would stay that way. Its roadmap included the things Snap had previously forked to get: <a href="https://valkey.io/blog/unlock-one-million-rps/">I/O threading</a>, <a href="https://github.com/valkey-io/valkey/issues/2083">dual-channel replication</a>, and a path toward features Snap wanted. And every major cloud provider was committing serious engineering effort to it.</p>
<p>The KeyDB story also got more complicated from the inside. In January 2025, John Sully, KeyDB’s original creator, <a href="https://github.com/Snapchat/KeyDB/issues/895">left Snap</a>. His parting note on the KeyDB repository said it plainly:</p>
<p>“When we made KeyDB we wanted to prove that caches should have great performance and I think we succeeded. Now there are many options, including Valkey which is fully open source and based on my testing has matched KeyDB’s performance. I’m not sure what Snap will do with the project, but I think that development effort should move to Valkey moving forward as they have clear momentum and are the most up to date.”</p>
<p>When the person who started the fork tells you the fork is done, the fork is done.</p>
<h2 id="the-secret-migration-back"><strong>The secret migration back</strong></h2>
<p>Snap runs caching at a scale where you can’t just swap a binary. The migration had to be invisible to application teams, comparable in cost, and safe across radically different workload types. Ovais walked through the major decisions that made their migration as easy as possible.</p>
<h3 id="abstraction-layers-are-key-to-managing-workloads-at-scale"><strong>Abstraction layers are key to managing workloads at scale</strong></h3>
<p>Snap had built a storage abstraction with a RESP proxy in front of every cluster. Applications never talked to KeyDB directly. They talked to the proxy, which spoke Redis wire protocol back to whatever was running behind it. That layer of indirection made this migration possible. Without it, every application team at Snap would have needed to know about the change. With it, nobody had to.</p>
<p>These layers let them migrate around 30 caches per week. By the time Ovais gave this talk, 70 to 80 percent of workloads were on Valkey.</p>
<h3 id="do-a-gap-analysis-before-changing-any-code"><strong>Do a gap analysis before changing any code</strong></h3>
<p>Snap did a feature-by-feature comparison between KeyDB and Valkey before touching anything in production. KeyDB’s multithreading and Valkey’s I/O threading work differently, so they benchmarked carefully to confirm comparable throughput. </p>
<p>Some KeyDB features were blockers and had to be ported to Valkey. Zone awareness was the first one Snap contributed. Replica MOVED behavior during upgrades was another. CPU throttling at high utilization was a third. </p>
<p>A hidden gap that wasn’t found until much later was with <code>MGET</code>. KeyDB supported it across slots, but Valkey does not. So after moving to Valkey, Snap had issues with command parsing pressure in large batching workloads. They quickly ported cross-slot <code>MGET</code> to their internal build, and are working with the core maintainers to get it added upstream.</p>
<h3 id="pick-a-stable-version-for-a-base-not-a-new-one"><strong>Pick a stable version for a base, not a new one</strong></h3>
<p>Snap started on Valkey 8.2 RC, ported the features they needed, and immediately ran into crashes at 9 to 10k QPS. The root cause was new TLS offloading work. They rolled back to 8.0.2, ported the necessary fixes onto that, and benchmarked from there. New releases need a baking period, and a migration is the wrong time to find out.</p>
<h3 id="categorize-and-prioritize-your-workloads"><strong>Categorize and prioritize your workloads</strong></h3>
<p>Snap divided their caches into three categories: CPU-bound, high-memory, and high-write-rate. Each category needed different validation. CPU-bound workloads were primarily a throughput question. High-memory workloads were really about replication buffer behavior during full syncs, because if the buffer fills before a snapshot completes, you enter a sync loop that never finishes. High-write workloads required tuning replica buffer sizes and primary write throttling, because Valkey’s dual-channel replication puts buffers on replicas rather than primaries. Inside each category, they went lowest-criticality first, highest-criticality last.</p>
<h2 id="lessons-from-going-full-circle"><strong>Lessons from going full circle</strong></h2>
<p>The fork was the right call in 2019. Redis was not going to go multithreaded, and the workloads Snap was running needed it. KeyDB was a solid piece of engineering that pushed the ceiling on what a single Redis-compatible node could do.</p>
<p>The migration back was the right call in 2025 because the conditions that justified the fork had changed. The upstream that resisted features they needed was no longer the upstream they cared about. Valkey’s governance was open. Its roadmap included the work Snap had previously done alone. And every additional year on a Redis 6.2 build was another year of compounding distance from where the ecosystem was going.</p>
<p>Forks are leverage. They are also debt. Be honest with yourself about which one you are accumulating at any given moment. Snap was. They forked when forking gave them speed, and they came back when the fork started to cost more than it earned.</p>
<p>I don’t want you to take away from this that forking is bad. Sometimes it’s the right thing to do. The decision to fork is not permanent, and treating it like it is permanent is how you end up running a five-year-old codebase while your competitors are shipping on a roadmap you helped fund.</p>
<p>When the world moves, move with it.</p>
<p>Happy coding!</p>]]></content:encoded>
  </item>
</channel>
</rss>
