Featured image: Circuit-board components illustrate the hardware underlying computing. Contextual photograph; not a verified NVIDIA accelerator or Dynamo deployment. Photo: Umberto / Unsplash. Unsplash licence.
Two people can send the same AI model very different jobs. One uploads a long document and wants a brief answer. Another sends a short instruction and asks for pages of output. A fleet of identical chips does not make those jobs identical.
This background explainer examines a design called disaggregated inference through NVIDIA Dynamo, announced on 18 March 2025, and current official documentation. It is an explanation of an infrastructure trade-off, rather than a new October product launch. The original technical announcement.
Reading the prompt and writing the answer are different phases
Inference means running a trained model to produce a result. In the prefill phase, a language model processes the input. In decode, it generates subsequent output tokens, units of text that may be words or parts of words. NVIDIA describes prefill as generally compute-bound and decode as generally memory-bound. NVIDIA explains the two phases.
Compute-bound work is limited mainly by how quickly calculations can be performed. Memory-bound work is limited mainly by moving the required information fast enough. These descriptions identify a likely bottleneck; the actual behaviour depends on the model, batch and hardware.
An aggregated deployment uses a worker for both phases. A disaggregated deployment assigns separate pools to prompt processing and generation, letting them be sized and scaled independently. The official deployment guide.

Specialisation creates a handoff
Prefill produces a key-value cache, commonly shortened to KV cache: saved intermediate information that decode uses when generating text. When different workers perform the two phases, the cache must be transferred before generation can continue. Dynamo’s architecture includes an inference-transfer library, NIXL, to handle this movement. Cache transfer in the Dynamo architecture.
A useful analogy is preparing ingredients in one kitchen and finishing the meal in another. Specialisation can improve each operation, but the prepared material has to arrive in time. If transport dominates, the advantage disappears.
The deployment documentation makes this trade-off explicit. Disaggregation can help with long prompts, high concurrency or different scaling requirements. For small models, short prompts, low concurrency or a slow transfer fabric, keeping the work together can be simpler and faster. The guide’s conditions and limitations.
The router must consider work already completed
Dynamo’s disaggregated router coordinates compatible prefill and decode services. Its documentation also discusses topology-aware transfer when workers are separated across zones or racks. The physical location of the cache can therefore affect where a request should go next. Prefill and decode routing documentation.
That turns a queue into an infrastructure decision. Sending a request to the least busy worker can look sensible until the cost of moving or rebuilding its context is included. Reusing an existing cache can save work, while overloading the worker holding it can create a delay.
The useful balance depends on traffic. A system serving occasional short chats has different pressures from one processing long documents for many simultaneous users. A single throughput figure cannot describe both.
Agents make cache lifetime part of the problem
In a technical analysis dated 17 April 2026, NVIDIA describes agent workflows that alternate model calls with tools and then reuse growing context. A cache can be valuable during a pause, even while another request competes for memory. The article distinguishes implemented mechanisms from further work on shared retention and lifecycle awareness. Read the agentic-inference analysis.
The practical tension is between keeping useful context and releasing scarce memory. Retaining everything can reduce room for new work. Discarding everything can force repeated computation when an agent resumes. Scheduling needs information about what is likely to happen next, rather than treating every request as unrelated.
Evaluation should consequently include the time to the first token, the pace of subsequent generation and performance under realistic concurrent demand. A useful test also counts transfer and recomputation costs, instead of measuring only the busiest calculation stage.
This is why inference infrastructure is becoming a coordination problem. Faster accelerators still matter, but the result also depends on giving each phase suitable resources and keeping its intermediate state in the right place at the right time.


Leave a Reply