Sparse AI Models Still Need Dense Infrastructure

Home

Rows of server cabinets in an imgix data centre

In brief

Mixture-of-experts models activate only part of their parameters for each token. The unused parts still shape memory requirements, and routing creates new work.

Featured image: Server cabinets illustrate computing infrastructure. Contextual photograph; not a verified DeepSeek-V3 deployment, MoE benchmark or particular expert-routing configuration. Photo: imgix / Unsplash. Unsplash licence.

A very large AI model can use only a small part of its numerical parameters while processing each token. That makes its active computation smaller than its full size suggests. It does not make the rest of the model vanish.

Mixture-of-experts, or MoE, architecture is an important example. This technical explainer uses current developer documentation and the earlier DeepSeek-V3 report. It describes an infrastructure trade-off, rather than a new October performance benchmark or a claim that one hardware platform always wins.

The experts are parts of a neural network

An expert is a neural subnetwork. It is not a human specialist or necessarily a clearly named skill such as medicine or mathematics. A learned router selects a subset of experts for a token, and their outputs are combined as the computation continues.

A token is a unit of input or generated text, often a word fragment. Routing takes place within the model’s layers, rather than simply choosing one expert for an entire conversation. NVIDIA’s architecture description explains this selective activation and the need to balance work among experts.

The appeal is straightforward: a model can contain more parameters without activating all of them for every token. The relevant comparison must still account for output quality, other parts of the network and the work required to organise those selected calculations.

Components and microchips on a printed circuit board
Electronic components illustrate computing hardware. Contextual photograph; not an identified AI accelerator, MoE router or evidence of model performance. Photo: Umberto / Unsplash. Unsplash licence.

Total size and active size are different numbers

The DeepSeek-V3 technical report, first submitted on 27 December 2024 and revised on 18 February 2025, describes 671 billion total parameters with 37 billion activated per token. These are useful architectural numbers from that report; they are not evidence of the latest October model ranking.

A parameter is a learned numerical value. The active count helps describe the computation selected for a token. The total count helps describe how much model data the deployment must accommodate. Storage precision and other serving requirements influence the actual memory footprint.

It follows that a sparse model cannot be sized using its active count alone. A deployment still needs access to its complete set of weights. Whether those weights are distributed across accelerators or managed through another supported arrangement changes the implementation, not the existence of the data.

Routing can turn computation into a communication problem

When experts are spread across GPUs, tokens must reach the devices holding their selected experts. Results then return to be combined. Megatron Core’s MoE documentation describes this dispatch-and-combine process, including all-to-all communication: participating devices exchange data with one another.

The same documentation identifies memory, communication and compute efficiency as performance constraints. Fewer active calculations therefore do not guarantee a proportionate reduction in elapsed time. Data movement and coordination can occupy time that is absent from a headline parameter count.

This does not make sparsity ineffective. It explains why its benefit depends on the system surrounding it. A model architecture and a network topology interact; assessing either in isolation can miss the reason a particular workload slows down.

Popular experts can create uneven demand

A router may send unequal amounts of work to different experts. NVIDIA describes load balancing as important during training and inference. Without suitable handling, a device serving a popular expert can become a hotspot while others have spare capacity.

The useful question is how the deployed system handles that distribution under representative traffic. An evenly distributed synthetic workload can miss difficulties that arise when many requests route similarly. This is an evaluation concern, rather than a measured failure of a named product.

The meaningful result is a complete service

Our assessment is that comparisons should report model quality, memory use, throughput and response delay under disclosed conditions. Batch size, weight precision, network arrangement and the traffic mix help explain a result. A processor’s peak arithmetic rate cannot supply those missing details.

Training and serving also ask different questions. Training must learn the parameters; serving uses learned parameters to produce responses. Documentation covering one stage should not be presented as a measured improvement in the other.

MoE makes selective computation a practical design tool. Its wider lesson is that efficiency must be traced through the whole system: the weights available, the work selected and the data exchanged. Sparse arithmetic still needs infrastructure that can deliver it.

Join the discussion

Have a question or a different perspective? Share it below. Please keep comments respectful and relevant to the article.

Leave a Reply

Your email address will not be published. Required fields are marked *

FUTURETECHDOSE BRIEFING

Follow the technologies shaping what comes next.

Clear, source-led reporting across biotechnology, AI infrastructure, energy, robotics and emerging devices.

Latest reporting