Kolibri’s Open Weights Make Sovereign AI a Concrete Infrastructure Choice

Home

Rows of server cabinets in an imgix data centre

In brief

Aleph Alpha’s new German-English model can run on infrastructure its users control. Open weights expand that choice while leaving hardware and validation work.

Featured image: Server cabinets at an imgix facility illustrate computing infrastructure. Contextual photograph; not Verda or the Kolibri training cluster. Photo: imgix / Unsplash. Unsplash licence.

Control over an AI system becomes more concrete when its model can be downloaded and run on infrastructure chosen by the organization using it. Aleph Alpha’s Kolibri release creates that option for a new German-English language model.

The company released Kolibri on 3 October 2026 under the Apache 2.0 licence. Its published model card lists approximately 78 billion total parameters and 3.46 billion active parameters per token. The release is a working model and documentation, rather than a promise to release weights later.

Open weights expand where the model can run

Weights are the numerical values learned during training. Access to them allows a model to be deployed without making every request through its developer’s hosted service, subject to the licence and practical requirements.

The model card lists an approximately 78 GB memory footprint for its FP8 weights. FP8 is a compact numerical representation. Running the model also requires memory and computing resources beyond storing those weights, especially as requests become longer or more numerous.

Its mixture-of-experts design selectively activates parts of the model for each token, a unit of text processing. The active-parameter count describes computation on that route through the model; it is not the size of the whole model that must be stored.

That distinction matters for public discussion of efficient AI. A model with a relatively small active count can still demand substantial hardware. “Open” describes access and permissions, not a guarantee of inexpensive deployment.

Components and microchips on a printed circuit board
An electronic circuit board illustrates the physical hardware underlying software. Contextual photograph; not a Kolibri-specific or NVIDIA product. Photo: Umberto / Unsplash. Unsplash licence.

A long context window has a practical operating range

Kolibri’s documentation supports a maximum context of 1,048,576 tokens while recommending at most 262,144 for serving efficiency and complex tasks. Context is the text available to the model during a request.

The maximum supported length and the recommended operating range answer different questions. The first describes a capability ceiling; the second helps define the conditions under which the developer expects more practical use.

Our assessment is that organizations should evaluate the documents and task sizes they actually need. A large context figure alone cannot show that a system reliably finds the right evidence in every long file.

The technical report provides the developer’s architecture and evaluation account. Its results are evidence from the model’s maker, and comparisons should retain their stated test conditions. They should not be presented as independent proof of universal superiority.

Training depended on a physical and software stack

In a 5 October 2026 account, infrastructure provider Verda says it supplied a custom bare-metal fleet using NVIDIA Blackwell GPUs and InfiniBand networking in European data centres. It describes monitoring, maintenance and recovery work alongside the hardware.

InfiniBand is a high-speed network used to connect machines working together. Bare metal means access to physical servers rather than only a virtual machine abstraction. Neither term establishes the quality of the resulting model by itself.

Aleph Alpha’s 22 May 2026 explanation of its training pipeline describes linking data, configurations, training and evaluation through versioned workflows. That background predates the model release and helps explain the engineering needed to reproduce a run.

A pipeline preserves how a model was made. Hardware supplies the computation. Evaluations examine what the model does. These layers complement one another, but a success in one cannot establish success in all three.

Sovereignty has to be defined for a particular deployment

Running downloaded weights can give an organization more control over where requests and sensitive documents are processed. The organization still needs to select hardware, manage access, monitor the service and test its outputs.

European training infrastructure does not imply that every component in the supply chain originated in Europe. Verda’s hardware account explicitly names NVIDIA equipment. Sovereignty therefore needs a stated meaning, such as deployment control or control of data, rather than a single sweeping label.

Nor does a downloadable model automatically make the entire training process reproducible. That would also depend on access to the relevant data, code, configurations and computing resources.

The significant development is a new deployment option with inspectable documentation and available weights. Its practical value will emerge from testing on real workflows: accurate use of supplied documents, manageable serving requirements and reliable operation under the organization’s chosen controls.

Join the discussion

Have a question or a different perspective? Share it below. Please keep comments respectful and relevant to the article.

Leave a Reply

Your email address will not be published. Required fields are marked *

FUTURETECHDOSE BRIEFING

Follow the technologies shaping what comes next.

Clear, source-led reporting across biotechnology, AI infrastructure, energy, robotics and emerging devices.

Latest reporting