Kolibri release details customer control and AI Act measures
Aleph Alpha’s 5 October announcement describes a German and English model, customer-controlled deployment and documented data-governance measures.
Aleph Alpha announced Kolibri in Heidelberg on 5 October 2026, describing it as a model for mission-critical work. The company says it supports German and English and is designed to run on customer-controlled infrastructure. Its release also outlines data-governance measures relevant to copyright, data protection and EU AI Act requirements.
The announcement sets out the company’s position and reported product features; it does not establish a regulatory decision or independent finding. The company’s technical report is identified as the place where its measures are documented. The release says the model had been available since 3 October.
Kolibri’s release combines deployment claims with an account of training-data checks. The stated arrangement gives customers control over the infrastructure on which the model runs. Aleph Alpha also says design decisions and data provenance are documented. Data provenance means information about where training data came from and how it was handled.
Training data checks are described by the company
Aleph Alpha says its technical report details steps addressing copyright and data protection, alongside EU AI Act requirements. Its release says all training data was screened against a blocklist of more than 4.5 million URLs. The company says sources for that list included the European Commission’s Piracy Watch List. The announcement does not describe a regulator’s review, a legal determination, or an enforcement action.
The company further says each third-party dataset underwent assessment for licence terms, lawful sourcing and adherence to opt-outs before use. Those statements describe Aleph Alpha’s process, rather than an external audit outcome. The release does not name an authority that has approved Kolibri or ruled on the adequacy of these checks. Readers seeking the underlying detail would need the technical report referenced by the company.
Model design targets German-language work
Kolibri is described as a 78 billion-parameter model, with around 3 billion active per token. Aleph Alpha says it uses a mixture-of-experts architecture that activates selected parts during inference. German-language text accounts for around 23% of pre-training data, according to the company, which says Kolibri is trained to reason in German. A German-optimized tokenizer is intended to represent German text with fewer tokens.
The announcement also lists support for agentic workflows, retrieval-augmented generation (RAG) and native tool calling. In RAG, a model uses supplied documents when answering questions. Aleph Alpha says Kolibri is trained to abstain if those documents lack sufficient evidence, and provides controls for balancing reasoning effort against response speed. The release gives no independent results for these capabilities.
Aleph Alpha presents the intended use
CEO Ilhan Scheer said the model reflects domestic capability and customer choice, while Co-founder and Co-Chief Research Officer Samuel Weinbach described its German-focused design. Their comments are the company’s response and rationale for the launch, not an authority’s assessment. Aleph Alpha says the intended settings are public administration and industry, and that Kolibri was developed and trained in Europe.
The announcement leaves open how the stated data checks are documented in full and how the model performs in deployments. It names no customers and reports no independent evaluations. No regulator’s decision or legal status beyond the company’s account of its measures is provided. The technical report and subsequent deployment evidence are the next materials that could clarify those points.
Sources
- Aleph Alpha releases Kolibri — Aleph Alpha aleph-alpha.com