How Sovereign Is Kolibri?

Aleph Alpha’s German foundation model is strongly sovereign in production and deployment, but not across the full technological stack

Kolibri demonstrates substantial European sovereignty over the production and deployment of a foundation model, but not full-stack technological autonomy. Aleph Alpha controls the architecture, training pipeline, post-training, evaluation and release process, while customers can retain and self-host the weights; critical dependencies remain in NVIDIA accelerators and the semiconductor stack, external teacher models, incomplete public reproducibility and future corporate governance. The article tests that conclusion against Kolibri’s Model Factory, its distance from the global capability frontier, Mistral Large 4, and the provenance of its synthetic training data.
EU digital sovereignty
machine learning
report
🇬🇧
Author

Antonio Montano

Published

October 5, 2026

Modified

October 7, 2026

Abstract

How sovereign is Kolibri? Substantially sovereign at the model-production and deployment layers, but not across the complete technological stack. Aleph Alpha controls the architecture, tokenizer, data-transformation pipeline, training recipe, post-training, evaluation and release process that produced Kolibri, while the resulting FP8 and BF16 weights can be retained and self-hosted without dependence on an Aleph Alpha inference service. The strongest sovereignty claim supported by the evidence is therefore not autarky, but control over the intellectual and operational system that turns data, compute and research decisions into a deployable foundation model.

That system is unusually well documented. Aleph Alpha’s Savanna Model Factory represents model development as version-controlled software, links datasets, configurations, checkpoints and evaluations through persistent lineage, automates experiments and recovery across large GPU clusters, and supports repeated model generations rather than a single training run. Kolibri itself is a 78.1-billion-parameter sparse Mixture-of-Experts model with 3.46 billion active parameters per token, trained through nearly 24 trillion base-model token presentations and followed by large-scale supervised fine-tuning and reinforcement learning. This constitutes strong evidence of European producer sovereignty and, because the weights are downloadable, substantial deployment sovereignty.

The limits are equally concrete. Kolibri’s demonstrated production system depends on NVIDIA B200 and B300 accelerators, high-bandwidth memory, networking and an international semiconductor and software ecosystem for which no equivalent European substitute is demonstrated. Its training-information supply chain is also not exclusively European: Aleph Alpha documents the use of GLM, Qwen, Gemma, Mistral and other external models to generate, regenerate, score or judge parts of the training data. The release is open-weight rather than a complete publicly reproducible production stack, and the announced Cohere transaction leaves future organizational control partly unresolved. Kolibri therefore does not establish full-stack or fully independent reproductive sovereignty.

The appendices sharpen that conclusion in two directions. A cross-benchmark calibration against Artificial Analysis places Kolibri heuristically in roughly the 20–26 Intelligence Index tier, well below the current global frontier overall, although the gap is much smaller on some mathematical, coding and long-context evaluations and much larger on broad knowledge reliability and long-horizon agentic execution. The October 2026 release of Mistral Large 4, independently scored at 38 and trained on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, shows that European model-development and infrastructure capability can move substantially closer to the frontier. At the same time, Kolibri’s disclosed dependence on external teacher models shows that sovereignty must include training-information provenance and substitutability as well as compute location and ownership.

The resulting judgment is therefore layered rather than binary: Kolibri demonstrates strong European model-production sovereignty and strong deployment sovereignty; partial reproductive and infrastructure sovereignty; and weak sovereignty at the accelerator and semiconductor layers. Europe has demonstrated that it can build, operate and retain serious foundation-model capability. It has not yet demonstrated that the entire chain required to reproduce the next generation can survive the loss of critical non-European technology, upstream models or organizational dependencies.

Kolibri demonstrates substantial European sovereignty over the production and deployment of a foundation model, but not full-stack technological autonomy. Aleph Alpha controls the architecture, training pipeline, post-training, evaluation and release process, while customers can retain and self-host the weights; critical dependencies remain in NVIDIA accelerators and the semiconductor stack, external teacher models, incomplete public reproducibility and future corporate governance. The article tests that conclusion against Kolibri’s Model Factory, its distance from the global capability frontier, Mistral Large 4, and the provenance of its synthetic training data.

The question is not whether Kolibri is European

Aleph Alpha released Kolibri on 3 October 2026 as a sovereign open-weight model. The company defines sovereignty along two axes: how the model is built and how control transfers to the customer. In the first dimension, it claims control and traceability from data ingestion through pre-training, post-training and evaluation; in the second, it points to downloadable weights and deployment freedom.1

That framing is useful because it shifts the discussion away from nationality as a label. The technically relevant question is not whether a model was announced by a German company, whether a datacentre is physically in Europe, or whether an API contract is governed by European law. It is how much of the dependency chain a European actor can actually inspect, operate, modify, retain and reproduce without discretionary permission from an external supplier.

Kolibri reaches further down that chain than a European fine-tune of a foreign base model. Aleph Alpha’s documentation shows internal work on corpus construction, tokenizer design, model architecture, MoE routing, optimizer selection, scaling studies, distributed training, long-context adaptation, supervised fine-tuning, reinforcement learning, grounding, evaluation and inference engineering.2 The same documentation describes the production system that connects these activities: the Model Factory called Savanna.

That distinction is the central thesis of this article. The strategically important result is not only that Aleph Alpha produced one 78-billion-parameter model. It is that it documents an organization and a software system that produced two large model generations three months apart, while changing almost every stage of the recipe between them.3

The sovereignty claim should nevertheless be scoped narrowly. Aleph Alpha’s full supply-chain integrity language is well supported if it means internal traceability across the model-development process. It is not evidence of European ownership of the NVIDIA accelerator stack, the semiconductor fabs and HBM supply chain behind those accelerators, or every upstream model used to generate synthetic training data. The rest of the article therefore separates model-production sovereignty from full-stack industrial sovereignty.

What Kolibri actually is

Kolibri is a decoder-only causal Transformer with a sparse Mixture-of-Experts architecture. Its total parameter count is 78,103,074,560, while 3,457,573,120 parameters are active for each token, excluding embeddings. The architecture contains 50 Transformer blocks, all using MoE feed-forward layers. Forty blocks use sliding-window attention over the previous 512 tokens; every fifth block uses full causal attention. Attention is grouped-query attention with 48 query heads, four key-value heads and head dimension 128. Each MoE layer contains 384 routed experts, activates six of them per token, and adds one shared expert.45

Property Kolibri
Developer Aleph Alpha Research GmbH
Provider Aleph Alpha GmbH
Architecture Decoder-only causal Transformer, sparse MoE
Total parameters 78.103B
Active parameters per token 3.458B
Transformer blocks 50
Sliding-window / full-attention blocks 40 / 10
Sliding window 512 preceding tokens
Model width 2,560
Query / KV heads 48 / 4
Head dimension 128
Routed experts per layer 384
Routed experts per token 6
Shared experts 1
Expert hidden dimension 512
Vocabulary 128,000
Tokenizer UniBPE
Native trained context 262,144 tokens
Validated extended context up to 1,048,576 tokens
Primary languages German and English
Released checkpoints FP8 and BF16
FP8 weight footprint about 78 GB
License on published weights/configuration Apache 2.0
Pre-training hardware 768 NVIDIA B200 GPUs
Pre-training compute approximately 6.4\times10^{23} FLOPs
Table 1: Core Kolibri specification from the official model card and technical report.

The architectural choice is explicitly cost-driven. Aleph Alpha tested larger sparse models around 123B total parameters and found continued quality gains, but the serving penalty at long context was severe. In the company’s comparison, a 123B candidate could fit only three concurrent 256k-token requests on two H100s, whereas the selected approximately 78B design could fit eighteen and decode 28% faster.6

This makes the sparse design part of the sovereignty argument rather than merely a research preference. A downloadable model that requires provider-scale infrastructure to run still transfers only limited operational autonomy. Kolibri was instead optimized so that a customer can serve it on a small number of high-end accelerators.

Figure 1: Average English and German benchmark score against decoded text per second and GPU. Source: Aleph Alpha technical report, Figure 1.

Figure Figure 1 shows the metric Aleph Alpha uses to justify that design point: quality against decoded bytes of text per second per GPU, rather than tokens per second. That choice matters for a bilingual model because tokenizer efficiency differs substantially across languages. The figure is still a developer-run comparison under Aleph Alpha’s serving setup, not an independent universal ranking, but it directly connects model architecture to deployment cost.

The real achievement is the Model Factory

The most important engineering object disclosed around Kolibri is Savanna, Aleph Alpha’s Model Factory. In the company’s terminology, Savanna implements Model Training as Code: the complete training workflow is represented as version-controlled software rather than as a set of manually coordinated scripts, filesystem paths and human hand-offs.7

Aleph Alpha describes three properties of this approach. Composability comes from representing pipeline stages as functions with typed inputs and outputs. Consensus comes from version control, where the main branch represents the current agreed training recipe. Provenance comes from commits, comments and artefact lineage, which allow a past run to be traced back to the code, configuration, datasets, tokenizer and checkpoints that produced it.8

The technical report is more concrete. Savanna owns orchestration of pre-training and post-training runs and their resumes, ablations and sweeps, checkpointing and model merging, evaluation, lineage, abstractions over clusters and storage, artefact retention and declarative model releases.9 A run is described declaratively, with versions of the dataset, tokenizer, architecture and evaluation strategy pinned as inputs.

%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%

flowchart TB

    %% ============================================================
    %% 1. DATA PRODUCTION
    %% ============================================================

    subgraph DATA["1 · Data production"]

        direction TB

        %% --------------------------------------------------------
        %% PRE-TRAINING DATA
        %% --------------------------------------------------------

        subgraph PREDATA["Pre-training data"]

            direction TB

            CC["Raw Common Crawl"]
            OPEN["Open / curated datasets<br/>web · documents · code · STEM · QA · legal · public sector"]

            EXTRACT["Resiliparse extraction"]

            HEUR["Language-agnostic heuristic filtering<br/>+ language identification"]

            LANG["English / German branch filtering<br/>language-specific thresholds"]

            DEDUP["Deduplication<br/>exact · fuzzy MinHash/LSH · substring"]

            QJ["Qwen3-32B<br/>judge on 1M English CC sample"]

            DISTILL["Distilled quality classifiers<br/>Luxical + fastText"]

            GEM["Gemma-4-26B-A4B<br/>English synthetic rephrasing"]

            MIS["Mistral-Nemo-Instruct-2407<br/>German rephrasing"]

            SAFE["Cross-source filtering and remediation<br/>illegal · harmful · toxic · pirated content<br/>PII detection and redaction"]

            PREPOOL["Pre-training candidate pool<br/>78 datasets · ~29.3T materialised tokens"]

            PRESEARCH["Pre-training mixture search<br/>thousands of proxy runs + mixing law"]

            PREMIX["Production pre-training mix<br/>20T tokens"]

            CC --> EXTRACT
            EXTRACT --> HEUR
            HEUR --> LANG
            LANG --> DEDUP

            DEDUP -. sample .-> QJ
            QJ --> DISTILL
            DISTILL -. quality scores .-> PREPOOL

            DEDUP -. English seeds .-> GEM
            DEDUP -. German documents .-> MIS

            DEDUP --> SAFE
            OPEN --> SAFE
            GEM --> SAFE
            MIS --> SAFE

            SAFE --> PREPOOL
            PREPOOL --> PRESEARCH
            PRESEARCH --> PREMIX
        end


        %% --------------------------------------------------------
        %% MID-TRAINING / LONG-CONTEXT DATA
        %% --------------------------------------------------------

        subgraph LATEDATA["Mid-training and long-context data"]

            direction TB

            CURATED["Curated capability-dense pool<br/>reasoning · maths · code · instruction data"]

            DECON["Benchmark decontamination<br/>226 evaluation tasks"]

            MIDSEARCH["Mid-training mix ablations<br/>domain + language search"]

            MIDMIX["Production mid-training mix<br/>3.44T tokens"]

            NATLONG["Naturally long documents<br/>8k–256k<br/>synthetic samples removed"]

            LONGMIX["Long-context production mix<br/>≈200B tokens<br/>natural + reweighted mid-training data"]

            OPEN --> CURATED
            PREPOOL -. selected sources .-> CURATED

            CURATED --> DECON
            DECON --> MIDSEARCH
            MIDSEARCH --> MIDMIX

            PREPOOL --> NATLONG

            MIDMIX --> LONGMIX
            NATLONG --> LONGMIX
        end


        %% --------------------------------------------------------
        %% SFT DATA
        %% --------------------------------------------------------

        subgraph SFTDATAFACTORY["Supervised fine-tuning data"]

            direction TB

            SFTSEED["Open / seed SFT datasets<br/>+ in-house German, long-context,<br/>agentic, retrieval, safety data"]

            GLM["GLM-5.2 / GLM-5.3"]

            Q38["Qwen3.8-27B"]

            SFTGEN["Synthetic generation / regeneration<br/>reasoning · German · long context<br/>agentic · science · multi-turn"]

            SFTPROC["Audits · filtering · decontamination<br/>reasoning-effort labelling"]

            SFTPOOL["Final SFT data pool<br/>57 datasets · 22.0M samples · 335.8B tokens"]

            MERGEMIX["MergeMix search<br/>cluster specialists + candidate weightings"]

            SFTRUNS["C1 / C2 full candidate runs<br/>~268B packed tokens each"]

            SFTSOUP["C1 + C2 weight soup<br/>released SFT checkpoint"]

            SFTSEED --> SFTGEN
            GLM --> SFTGEN
            Q38 --> SFTGEN

            SFTSEED --> SFTPROC
            SFTGEN --> SFTPROC

            SFTPROC --> SFTPOOL
            SFTPOOL --> MERGEMIX
            MERGEMIX --> SFTRUNS
            SFTRUNS --> SFTSOUP
        end
    end


    %% ============================================================
    %% 2. DEVELOPMENT / EXPERIMENT CONTROL
    %% ============================================================

    subgraph DEV["2 · Model development and recipe control"]

        direction TB

        GIT["Git repository<br/>Savanna + shared training recipe"]

        MAIN["main branch<br/>collective current recipe"]

        BRANCH["Experiment / protected branches"]

        CFG["Declarative TOML run configuration<br/>architecture · datasets · tokenizer<br/>training · checkpointing · evaluation"]

        CI["Pull-request CI<br/>small end-to-end training tests"]

        NIGHT["Nightly end-to-end<br/>train + evaluate regression run"]

        REVIEW["Reviewed pull request"]

        EXP["Experiments / ablations / sweeps"]

        PROXY["Proxy-scale model studies<br/>architecture · optimisation · data"]

        SCALE["Systems scaling experiments<br/>FSDP · DP · EP · activation checkpointing"]

        POSTEXP["Post-training experiments<br/>SFT mixes · RL environments"]

        DECIDE["Evidence-backed recipe change"]

        GIT --> MAIN
        MAIN --> BRANCH
        BRANCH --> CFG
        BRANCH --> EXP

        CFG --> CI
        CI --> REVIEW
        REVIEW --> MAIN

        MAIN --> NIGHT

        EXP --> PROXY
        EXP --> SCALE
        EXP --> POSTEXP

        PROXY --> DECIDE
        SCALE --> DECIDE
        POSTEXP --> DECIDE
    end


    %% ============================================================
    %% 3. IMMUTABLE ARTEFACT / LINEAGE LAYER
    %% ============================================================

    subgraph ART["3 · Immutable artefact and lineage layer"]

        direction LR

        VD["Versioned datasets"]
        VT["Versioned tokenizers"]
        VC["Versioned configurations"]
        VCP["Input / predecessor checkpoints"]

        REG["Immutable artefact registry<br/>versions · retention · provenance"]

        VD --> REG
        VT --> REG
        VC --> REG
        VCP --> REG
    end

    PREMIX --> VD
    MIDMIX --> VD
    LONGMIX --> VD
    SFTPOOL --> VD

    MAIN --> VC


    %% ============================================================
    %% 4. SAVANNA CORE
    %% ============================================================

    subgraph SAV["4 · Savanna — Model Training as Code"]

        direction TB

        ORCH["Pipeline orchestration<br/>runs · resumes · sweeps · ablations<br/>merges · evaluations · releases"]

        CPA["Checkpoint abstraction<br/>storage · retention · versioning<br/>lineage · DCP / Hugging Face conversion"]

        BA["Benchmark abstraction<br/>deployment · inference settings<br/>backend · scoring · aggregation"]

        CA["Cluster abstraction<br/>GPU type · workload manager<br/>storage · cluster-specific details"]

        RELEASECFG["Declarative release definition"]

        ORCH --> CPA
        ORCH --> BA
        ORCH --> CA
        ORCH --> RELEASECFG
    end

    CFG --> ORCH
    REG --> ORCH


    %% ============================================================
    %% 5. EXECUTION INFRASTRUCTURE
    %% ============================================================

    subgraph INFRA["5 · Workflow, software and compute infrastructure"]

        direction TB

        FLYTE["Flyte<br/>durable workflows · retries · caching"]

        K8S["Kubernetes"]

        KUEUE["Kueue<br/>priority + topology-aware scheduling"]

        CLUSTERS["Multiple GPU clusters"]

        B200["Base-model infrastructure<br/>768 NVIDIA B200 · 96 × 8-GPU nodes"]

        B300["Main RL infrastructure<br/>256 NVIDIA B300 · 32 nodes"]

        STACK["ML software substrate<br/>PyTorch / TorchTitan fork<br/>DeepEP · FlashAttention · vLLM"]

        NET["GPU interconnect<br/>NVLink · InfiniBand · ConnectX-7"]

        OBS["Observability / diagnostics<br/>Grafana · logs · metrics<br/>DCGM · PyTorch profiling"]

        CA --> FLYTE
        FLYTE --> K8S
        K8S --> KUEUE
        KUEUE --> CLUSTERS

        CLUSTERS --> B200
        CLUSTERS --> B300

        B200 --> NET
        B300 --> NET

        CLUSTERS --> OBS
    end


    %% ============================================================
    %% 6. BASE-MODEL TRAINING
    %% ============================================================

    subgraph BASE["6 · Base-model production"]

        direction TB

        PRE["Pre-training<br/>20T · 16k context<br/>EP8 · FSDP16 · DP6"]

        MID["Mid-training<br/>3.44T · 64k context<br/>FSDP128 · DP6"]

        LONG["Long-context extension<br/>≈200B · 256k context<br/>FSDP128 · DP6"]

        WSM["Warmup-Stable-Merge<br/>average final 20 checkpoints"]

        BASEMODEL["Kolibri Base"]

        PRE ==> MID
        MID ==> LONG
        LONG ==> WSM
        WSM ==> BASEMODEL
    end

    PREMIX --> PRE
    MIDMIX --> MID
    LONGMIX --> LONG

    B200 --> PRE
    B200 --> MID
    B200 --> LONG

    STACK --> PRE
    STACK --> MID
    STACK --> LONG


    %% ============================================================
    %% 7. SFT
    %% ============================================================

    BASEMODEL --> SFTRUNS

    CLUSTERS --> SFTRUNS
    STACK --> SFTRUNS


    %% ============================================================
    %% 8. REINFORCEMENT LEARNING
    %% ============================================================

    subgraph RLSTAGE["7 · Reinforcement learning"]

        direction TB

        RLTASKS["38 RL environments<br/>1.2M+ internally curated tasks"]

        POLICY["Policy inference<br/>144 B300 GPUs<br/>vLLM · FP8"]

        RLJUDGE["Frozen RL judge<br/>Qwen3.5-122B-A10B<br/>16 B300 GPUs"]

        WORKERS["Rollout workers<br/>curriculum + environment stepping"]

        SANDBOX["Apptainer sandboxes<br/>persistent task state<br/>resource limits · allowlisted egress"]

        REPLAY["Replay buffer<br/>completed rollout groups"]

        TRAINER["RL trainer<br/>96 B300 GPUs · FSDP<br/>1000 optimiser steps"]

        FINAL["Released post-trained Kolibri"]

        RLTASKS --> WORKERS

        WORKERS --> POLICY
        WORKERS --> RLJUDGE
        WORKERS --> SANDBOX

        POLICY --> REPLAY
        RLJUDGE --> REPLAY
        SANDBOX --> WORKERS

        REPLAY --> TRAINER

        TRAINER -. FP8 weight push each step via RDMA .-> POLICY

        TRAINER ==> FINAL
    end

    SFTSOUP --> TRAINER

    B300 --> POLICY
    B300 --> RLJUDGE
    B300 --> TRAINER

    STACK --> POLICY
    STACK --> TRAINER

    GLM -. "GLM-5.2 also generates<br/>some RL long-context questions" .-> RLTASKS


    %% ============================================================
    %% 9. CHECKPOINT / EVALUATION LOOP
    %% ============================================================

    subgraph EVAL["8 · Continuous checkpoint and evaluation loop"]

        direction TB

        ICP["Intermediate checkpoints"]

        CPMG["Configured checkpoint-window merge"]

        STATIC["Static benchmarks<br/>knowledge · maths · code<br/>German · English"]

        AGENTBENCH["Agentic benchmarks<br/>tool use · retrieval · environments"]

        INDUSTRY["Industry / customer proxies<br/>semiconductors · public sector<br/>aerospace · automotive · drives"]

        GROUND["Grounding / abstention<br/>hallucination evaluations"]

        LONGEVAL["Long-context evaluations"]

        SCORES["Scores · metrics · run comparison"]

        ICP --> CPA
        CPA --> CPMG

        CPMG --> STATIC
        CPMG --> AGENTBENCH
        CPMG --> INDUSTRY
        CPMG --> GROUND
        CPMG --> LONGEVAL

        BA --> STATIC
        BA --> AGENTBENCH
        BA --> INDUSTRY
        BA --> GROUND
        BA --> LONGEVAL

        STATIC --> SCORES
        AGENTBENCH --> SCORES
        INDUSTRY --> SCORES
        GROUND --> SCORES
        LONGEVAL --> SCORES
    end

    PRE --> ICP
    MID --> ICP
    LONG --> ICP
    SFTRUNS --> ICP
    TRAINER --> ICP

    CPA -. lineage .-> REG
    SCORES -. evaluation lineage .-> REG

    SCORES -. stop / reject unpromising runs .-> ORCH


    %% ============================================================
    %% 10. ORGANISATIONAL LEARNING
    %% ============================================================

    subgraph LEARN["9 · Organizational learning loop"]

        direction TB

        HUMANS["Researchers / engineers"]

        AGENTS["Software agents<br/>experiment review · failed-run triage<br/>GPU-utilisation optimisation"]

        ANALYSE["Analyse metrics · failures<br/>profiles · benchmark regressions"]

        CHANGE["Propose recipe / code change"]

        ISSUES["Repository issues<br/>known limitations · XXX markers"]

        HUMANS --> ANALYSE
        AGENTS --> ANALYSE

        ANALYSE --> CHANGE
        ANALYSE --> ISSUES

        CHANGE --> REVIEW
        ISSUES --> GIT
    end

    SCORES -.-> ANALYSE
    OBS -.-> ANALYSE
    NIGHT -. regression results .-> ANALYSE

    AGENTS -.-> GIT
    AGENTS -.-> EXP
    AGENTS -.-> OBS
    AGENTS -.-> SCORES

    DECIDE --> CHANGE


    %% ============================================================
    %% 11. RELEASE AND DEPLOYMENT
    %% ============================================================

    subgraph REL["10 · Release and customer deployment"]

        direction TB

        SELECT["Selected final checkpoint"]

        FP8["FP8 export"]
        BF16["BF16 export"]

        HF["Hugging Face release<br/>weights + configuration<br/>Apache 2.0"]

        DOC["Public documentation<br/>model card · technical report<br/>training-content summary"]

        SERVE["Self-hosted inference<br/>aleph-alpha-inference + vLLM"]

        API["OpenAI-compatible API"]

        SELECT --> FP8
        SELECT --> BF16

        FP8 --> HF
        BF16 --> HF

        HF --> SERVE
        SERVE --> API

        SELECT -. documented by .-> DOC
    end

    FINAL --> SELECT
    RELEASECFG --> SELECT

    REG -. provenance / configuration .-> DOC
    SCORES -. evaluation results .-> DOC
Figure 2: Aleph Alpha’s documented Kolibri Model Factory and production pipeline. The diagram separates pre-training data construction, later-stage decontaminated data, SFT synthetic-data generation, Savanna orchestration, distributed training, RL execution, evaluation feedback, lineage and release.

Figure Figure 2 summarizes the documented relationship rather than adding a new architectural claim. The key point is that Savanna is not described as a trainer alone. It is the integration layer that binds training code, artefacts, compute and evaluation into a reproducible internal workflow.

The pipeline has explicit core abstractions. Its Checkpoint abstraction manages storage, retention, versioning, lineage and format conversion, and connects a checkpoint to the configuration, commit, datasets and evaluation results that produced it. A Benchmark abstraction encapsulates the evaluation backend, inference settings, deployment, scoring and aggregation for everything from static question sets to multi-turn tool-using environments. A Cluster abstraction hides differences among GPU types, workload managers and storage systems so that the logical recipe can be dispatched to different clusters.10

This is internal portability, not proof of hardware independence. The same recipe can be expressed against multiple cluster backends; that does not establish equivalent performance on a non-NVIDIA accelerator stack.

Continuous integration for foundation-model training

The pipeline is operated with software-engineering controls more commonly associated with production application code than with one-off research runs.

Savanna lives in GitHub, and its continuous-integration path is the entry point for training. Aleph Alpha says pull requests can trigger a small-scale end-to-end training run that completes in under five minutes. A larger end-to-end model is trained and evaluated nightly to detect semantic regressions in training and evaluation code. The team uses trunk-based development so changes reach main in small increments rather than accumulating in long-lived research branches.11

Non-code artefacts are also versioned. Aleph Alpha says model, data and tokenizer artefacts are stored immutably, with runs linking the referenced artefacts, logs, metrics, evaluation results and resulting checkpoints. The company’s engineering post describes an on-premises object store, Weights & Biases for artefact lineage, Flyte as the workflow engine, Kubernetes for the primary clusters and Grafana for operational visibility.12

The technical report adds Kueue for priority-based and topology-aware scheduling, specifically to keep jobs inside an InfiniBand domain where required.13 The point is not that these components are unique. Flyte, Kubernetes, Grafana and the surrounding open-source software ecosystem are widely available. The achievement is the integration of those components into a model-development process in which the training recipe, data lineage, checkpoints, evaluation and cluster execution are tied together strongly enough that a large run can be restarted, inspected and changed without reconstructing its state from human memory.

The production run is evidence of the pipeline, not just a claim about it

Kolibri Origin and Kolibri provide the strongest operational evidence that Savanna works at scale. Aleph Alpha says work on the pipeline began in January 2026. Kolibri Origin completed target-scale pre-training on 11 June; Kolibri completed pre-training on 11 September and was released on 3 October. In those three months, the model moved from 30.6B to 78.1B total parameters, from 7.51T to 20T pre-training tokens, from 65k to 256k maximum trained sequence length, from 128 to 384 routed experts and from full attention in every layer to the 4:1 sliding-window/full-attention design.14

The production pre-training run lasted 21 days on 768 B200 GPUs. Aleph Alpha reports 38 unplanned interruptions, approximately one per 10,000 GPU-hours, caused by hardware faults or connection timeouts. The pipeline restarted training automatically on different nodes and resumed from a checkpoint no more than 250 optimizer steps behind, without requiring a person to reconstruct the run.15

The report gives a second operational example. Production pre-training ran from a protected branch derived from main; Savanna automatically evaluated a merge of the last four checkpoints every 2,000 steps. A mid-run change to logging and profiling frequency increased throughput by about 3.7%, while the run remained represented as one lineage-preserving training job despite restarts.16

Figure 3: Best checkpoint performance reached over time during the Kolibri Origin and Kolibri development cycles. Source: Aleph Alpha technical report, Figure 49.

Figure Figure 3 is therefore better interpreted as evidence about an engineering organization than as another model benchmark. Aleph Alpha’s stated objective is to shorten the time between identifying a failure and training a checkpoint that performs better on the corresponding evaluation.

Efficiency engineering is itself part of the capability

A separate Aleph Alpha engineering study gives unusually concrete detail on how the team scales pre-training jobs. It is important to distinguish this study from the final Kolibri production run: the scaling experiment used a 30B-total, 3B-active MoE rather than the released 78B model. Its relevance is that it exposes the methodology used by the efficiency team.17

The study starts at 8–16 B200 GPUs, where the team sweeps combinations of FSDP degree, data-parallel degree, activation checkpointing and local batch size. The small-scale sweep is used to eliminate configurations before scaling. PyTorch profiles then identify the boundary between compute-bound and communication-bound execution.

At 128 GPUs, selective activation checkpointing with FSDP128 reached 28.0k tokens/s/GPU, 37.0% model-FLOPS utilization and 93% HBM usage. The final 512-GPU configuration used FSDP128, DP4, selective activation checkpointing and local batch size 22, reaching 26.7k tokens/s/GPU and 35.3% MFU with 94% HBM usage. Per-GPU efficiency fell by only about 6% between the 16-GPU optimum and the 512-GPU result.18

This is a meaningful achievement, but it should not be overstated. Aleph Alpha explicitly declines to compare the 35.3% MFU number directly with unrelated published MoE runs because architecture, precision and recipe differences make such comparisons unreliable. The documented result is the near-linear scaling of this particular training setup, not proof that it is the industry’s most efficient MoE implementation.19

The final Kolibri production topology is different again: 768 B200 GPUs across 96 eight-GPU nodes, with expert parallelism of 8, FSDP of 16 and data parallelism of 6 during pre-training. Within a node the GPUs use fifth-generation NVLink; nodes are connected through InfiniBand with eight ConnectX-7 adapters per node.20

Taken together, the two sources establish a hard fact that matters for sovereignty: whatever the commercial ownership arrangement for the physical compute, Aleph Alpha is clearly not merely consuming a managed model-training API: its engineers operate the distributed training, profiling, parallelism, recovery and evaluation stack themselves. It has a team capable of profiling, parallelizing, debugging and operating large MoE training workloads at hundreds of current-generation accelerators.

The data factory starts far before the 20T-token run

The release announcement says the pipeline processed more than 200 trillion raw tokens to produce the final 20T-token pre-training run.21 The technical report describes an intermediate materialized pool of approximately 29.3T tokens across 78 datasets, from which the production mix consumes 20T.22

Those numbers describe different levels of the data pipeline and should not be collapsed. The 200T figure is the scale of raw material passing through extraction, filtering and curation; 29.3T is the materialized candidate pool; 20T is the final pre-training consumption.

The Common Crawl pipeline is similarly explicit. The report says Aleph Alpha extracts more than 308 billion documents, performs exact deduplication to obtain approximately 21.16 billion distinct documents, then applies fuzzy MinHash/LSH deduplication that removes a further 25.9% of near-duplicate documents. Substring deduplication removes repeated blocks such as navigation and templated page chrome, eliminating 21.3% of text bytes from the remaining corpus.23

The pre-training data pipeline also performs quality scoring, URL and content filtering and PII remediation. Benchmark decontamination is applied explicitly to later-stage training pools rather than to the complete pre-training pool. The model card states that high-syntax-constraint identifiers such as email and IP addresses are replaced with special tokens and that documents are dropped when more than 15% of their bytes are replaced during redaction.2425

About 24% of the final pre-training token presentations are synthetic in origin, including rephrases and translations.26 This is important for the sovereignty assessment because Aleph Alpha controls the generation and filtering process while not originating every upstream model used to create that material.

German is treated as a data-engineering problem, not a translation feature

Kolibri’s German capability is one of the clearest examples of an engineering decision that would be invisible in a generic European model label.

Aleph Alpha’s small-model ablations indicated that a German share around 20% was useful at a 20T-token pre-training horizon. The final production mix contains 21.3% German token presentations.27 Open German datasets were insufficient to reach the target, so Aleph Alpha built a language-specific Common Crawl pipeline and synthetic rephrasing process.

The company explicitly reports that English filtering rules cannot simply be applied to German. One example is mean word length: thresholds tuned on English can discard formal German administrative prose because German compounds are longer. Aleph Alpha therefore retuned language-specific filters rather than accepting the distribution imposed by an English-centric pipeline.28

The sources are not perfectly numerically consistent on every intermediate count. The launch article says the German web pipeline produced 1.3T unique organic tokens, while the detailed technical-report narrative says 0.8T unique tokens after the language-specific processing described there. Both sources agree on the larger operational point: the final German pool is about 2.4T unique tokens, roughly 80% curated or generated by Aleph Alpha and 20% from open datasets, and it is upsampled to about 4.3T German token presentations during the 20T pre-training run.2930 This article does not attempt to reconcile the differing 0.8T and 1.3T intermediate accounting boundaries.

The tokenizer is part of the same design. Kolibri uses a 128k-vocabulary tokenizer trained with UniBPE, which combines BPE’s bottom-up merge structure with a Unigram-loss criterion for selecting merges. In Aleph Alpha’s comparison, the tokenizer reaches about 4.90 UTF-8 bytes per token on German FineWeb-2, compared with 4.35 for GPT-5, while remaining close to the leading tokenizers on English.31

Figure 4: Tokenizer compression on German FineWeb-2 and English FineWeb. Higher bytes per token means more source text is represented by each token. Source: Aleph Alpha technical report, Figure 5.

Figure Figure 4 connects linguistic specialization to economics. More German bytes per token means fewer model tokens for the same document, which reduces prefill work, KV-cache pressure and token-metered serving cost.

German reasoning exposed a failure mode instead of hiding one

Aleph Alpha’s September research on German reasoning is relevant not because every result transfers directly to the final model, but because it shows how the post-training pipeline was used to create and test a capability that open datasets did not supply.

The team generated 795,731 German SFT conversations across general chat, tool use, mathematics, science and code; 83% of those conversations contained German reasoning traces.32 Because strong reasoning teachers tended to switch to English internally, the pipeline prefills the beginning of the teacher’s reasoning trace with a German opener. In the reported measurement, a system-prompt-only approach kept 67% of traces in German, while the prefilled-opener approach kept 97% in German.33

The experiment also reported a negative result. Introducing a small amount of German reasoning data caused some models to enter repetition loops and fail to produce a final answer. On German AIME 2026, one low-data run fell from a 70.2 baseline to 48.3; increasing the amount of domain-matched German math data recovered the score to 67.3, but did not fully eliminate the looping behavior.34

This is useful evidence about transparency in a technical sense. The publication does not present German reasoning as a solved binary capability. It documents the data-generation method, the measured language-consistency gain, the failure mode, and the fact that domain-specific German data mattered more than German data in general.

Pre-training, mid-training and long-context training are separate curricula

Kolibri Base is not produced by one homogeneous 24T-token run.

Stage Token budget Context length Global token batch
Pre-training 20T 16,384 about 75.5M tokens/step
Mid-training 3.44T 65,536 about 100.7M tokens/step
Long-context extension 201B 262,144 about 201.3M tokens/step
Table 2: Base-model training stages documented for Kolibri.

Pre-training uses the broad data mixture. Mid-training shifts toward capability-dense material such as reasoning, mathematics, code and agentic data. The long-context phase mixes long documents with high-quality mid-training data rather than training only on long documents, because Aleph Alpha reports that long-document-only adaptation can erode previously acquired capabilities.35

The optimizer is split by parameter group. Aleph Alpha uses spectral Muon for attention and expert matrix parameters, with Adam or AdamW for routers, embeddings, normalization parameters and the language-model head. The report specifies a 100B-token warm-up followed by a constant learning rate across the remaining base-model stages, with optimizer state carried forward rather than reset.36

The final base checkpoint uses a Warmup-Stable-Merge recipe: rather than adding a conventional learning-rate cooldown, the final weights are obtained by averaging the last 20 checkpoints from the long-context stage.37

A separate Aleph Alpha study on a 30B-total, 3B-active model helps explain why the company treats stage boundaries cautiously. Three pre-training checkpoints that ranked differently immediately after pre-training changed order after the same downstream mid-training, long-context and SFT pipeline; the study concludes that checkpoint selection should account for the training that remains.38 That result should not be treated as a proof about every Kolibri checkpoint, but it documents the model-development discipline behind the factory: early benchmark scores are not assumed to be sufficient selection criteria for a multi-stage training system.

Mixture search is itself industrialized

The pre-training mix was not chosen by hand once and then frozen. The technical report says the data team trained thousands of small proxy models, each on a different mixture, and fitted a mixing law to their scores. The official model card further specifies that the pre-training search used 30M-parameter dense proxies trained on 3B tokens.3940

The same pattern appears later in the pipeline. Mid-training candidates are compared with proxy models, and Savanna can compose a short SFT stage onto each candidate before evaluation so that the team receives a downstream signal rather than judging the candidate solely at the point where the intervention occurred.41

For SFT, Aleph Alpha adapted MergeMix to twenty data clusters. The model card says it trained one specialist per cluster, merged 69 candidate weightings drawn around the hand-tuned mix, evaluated them across sixteen benchmarks in seven capability groups, trained eight candidate mixtures at proxy scale, and moved the four best to target-scale tests.42

This is the operational meaning of Model Factory: not just automation of a final known recipe, but a system for repeatedly searching the recipe itself.

Post-training is another large production pipeline

Kolibri post-training has two principal stages: SFT followed by reinforcement learning. The release article says Aleph Alpha generated 174B synthetic SFT tokens, filtered them and combined them with permissively licensed data into a 268B-token training mixture.43 A full SFT run uses 4,000 optimization steps at 67.1M packed tokens per step, or about 268.4B token positions. The final SFT checkpoint is a model soup of two checkpoints trained for the same horizon on different data mixtures.44 The technical report’s statement that SFT spans about 537B training-token presentations is therefore consistent with two full approximately 268B candidate runs contributing to the final soup.45

The data provenance is important. The technical report explicitly states that the main models used to generate SFT data and regenerate parts of open datasets are GLM-5.2, GLM-5.3 and Qwen3.8-27B.46 Different teacher models are used for different generation and judging tasks because Aleph Alpha finds that they have different capabilities, values and biases.

That means Kolibri is not a derivative fine-tune of GLM or Qwen. Its base model was trained by Aleph Alpha from its own architecture and data pipeline. But the final model has been post-trained on synthetic examples produced by non-European foundation models. This is an upstream training-information dependency, not a runtime API dependency.

The model card acknowledges the associated political-bias risk and says Aleph Alpha filters SFT data for political bias and adds dedicated alignment data grounded in curated material on politically sensitive topics.47 A separate Aleph Alpha study on Chinese model alignment reports that the company adopted screening, dedicated alignment data and explicit evaluation for this purpose.48

For reinforcement learning, Aleph Alpha reports more than 1.2 million internally curated tasks across reasoning, tool use, instruction following, code, retrieval and other domains.49 The main RL run lasts 1,000 optimizer steps at sequence lengths up to 256k. Rollouts are produced through vLLM while training proceeds asynchronously, with updated weights synchronized to inference without waiting for all ongoing generations to finish. Aleph Alpha also uses quantization-aware training so the post-trained model can be served efficiently at low precision.50

This is another important sovereignty boundary. Aleph Alpha owns the RL environments, training orchestration and final weights, but some SFT information is generated by external model families and the runtime stack still contains globally developed software.

Grounding and abstention are deliberately trained behaviors

One distinctive part of Kolibri’s post-training programme is the attempt to make unsupported answers visible through abstention. Aleph Alpha trains with conventional abstention data and with a Merlin-Arthur protocol. In the company’s description, one player constructs a context that supports the correct answer, while another removes relevant evidence; the model is trained to answer in the former case and abstain in the latter.51

Figure 5: Grounding and hallucination evaluations for Kolibri, Kolibri Origin and selected open models. The M/A grounding spoke uses a 0–0.5 scale rather than a 0–100 percentage scale. Source: Aleph Alpha technical report, Figure 48.

Figure Figure 5 shows the developer’s own measurements. The AA-Omniscience non-hallucination rate rises from about 15% for Kolibri Origin to 44% for Kolibri; RGB negative-condition abstention rises from about 74% to 86%; and the Merlin-Arthur grounding score reaches approximately 0.234 where Origin scores zero.52

The M/A number is not an accuracy percentage. Aleph Alpha describes it as a lower-bound certificate of how much a correct answer depends on the supplied document. Its scale is different from the other axes in the radar plot and should not be read visually as if every spoke shared the same unit.

Customer proxies show the hill-climbing loop, not independent validation

Aleph Alpha also evaluates successive RL checkpoints against internal customer proxies for semiconductors, the German public sector, aerospace, automotive suppliers and industrial drive technology. The report says these evaluations use production-like prompts, tools and separate corpora, and that their documents and questions do not overlap with the training data.53

Figure 6: Evolution of agentic retrieval and internal Industry RAG customer-proxy evaluations during September 2026 reinforcement learning. Source: Aleph Alpha technical report, Figure 44.

Figure Figure 6 shows large improvements across the RL programme, including the semiconductor proxy moving from 35.3 to 80.4, the German public-sector proxy from 54.0 to 75.0 and aerospace from 14.1 to 58.9.54

These are not independent customer benchmarks. Aleph Alpha controls the evaluation construction and grading. Their value is different: they show the feedback mechanism of the Model Factory, where a domain failure can be represented as an evaluation and training environment and then tracked across checkpoints.

The public benchmark record also deserves a bounded interpretation. Kolibri performs strongly among the sparse models in Aleph Alpha’s comparison, especially in mathematics and code, but the dense Qwen3.8-27B scores higher on the report’s overall English and German aggregates.55 Kolibri is therefore evidence of an efficient and specialized European model, not evidence that Europe has already matched the global frontier on all capability dimensions.

Figure 7: Selected post-training benchmarks across agentic tasks, industry, mathematics/code and general knowledge. Source: Aleph Alpha technical report, Figure 2.

Transparency: unusually detailed, but not full external reproducibility

Aleph Alpha makes a deliberate transparency claim around Kolibri, and the public record supports part of it strongly.

The release includes a long technical report, a detailed model card, a public training-content summary, exact architecture parameters, training-stage token counts, optimizer details, hardware topology, major post-training hyperparameters, benchmark tables, energy estimates and deployment instructions. The model card itself says it was auto-generated by Savanna at a specific Git commit, which ties public release documentation back to the internal production system.56

Aleph Alpha has also signed the EU General-Purpose AI Code of Practice and, in August 2026, the Code of Practice on Transparency of AI-generated Content. In its own public statement, the company frames machine-readable provenance and transparency as part of its sovereignty strategy.57

That policy commitment should not be confused with a feature already present in every Kolibri output. The model card explicitly states that content generated by the model is not explicitly detectable at this point and that downstream systems must mitigate the risk of outputs being mistaken for human content.58

The important distinction is therefore between transparency about the model and its production process and machine-verifiable provenance of every generated output. The first is unusually extensive for a commercial model. The second is not claimed as a completed Kolibri capability in the model card.

Open weights are not the same as open source

Kolibri is downloadable in FP8 and BF16 form. The FP8 model has a memory footprint of roughly 78 GB, and Aleph Alpha documents serving configurations as small as one H200, B200 or B300, or two H100 SXM5 GPUs.59 The documented serving path uses Aleph Alpha’s inference package as a vLLM plugin and exposes an OpenAI-compatible API.

That transfers substantial control to the deployer. Once an organization has lawfully obtained and retained the weights, it can run that model version without an Aleph Alpha-hosted inference service. Sensitive prompts and retrieved documents can remain inside the organization’s own infrastructure.

The release should nevertheless be called open-weight, not fully open-source or fully reproducible. The model card is explicit: Apache 2.0 applies to the weights and configuration files published in the repository. Other artefacts not present in the repository are excluded from that grant, and Aleph Alpha states that the license does not extend to underlying code, model architecture, parameter settings or training methods. The company retains rights to those artefacts and methods.60

This creates an unusual but coherent disclosure model. Aleph Alpha publishes technical descriptions of many architectural parameters and training methods while not granting the complete production stack under the same open license. Savanna itself is described in detail but is not released as the public reproduction package for Kolibri. The raw and synthetic training datasets are not published as a complete reconstructable corpus.

The result is strong artefact transparency and deployment freedom, but incomplete third-party reproductive sovereignty.

Layer Publicly available Not publicly transferred as part of the Kolibri Apache release
Model weights FP8 and BF16 checkpoints —
Model configuration published —
Architecture description extensively documented architecture rights excluded from Apache grant
Training recipe extensively documented in report/model card complete production implementation not released
Inference path documented vLLM-based serving package complete internal serving/operations stack not implied
Training data content summary, categories, curation methods complete reconstructable corpus not published
Model Factory architecture and workflow described Savanna production code not released with the weights
Experiment lineage internally maintained and partly exposed through report/model card full internal experiment history not public
Table 3: Kolibri’s transparency/open-weight boundary.

This distinction matters for sovereignty because two different actors receive two different forms of control. Aleph Alpha retains the capability to build successor models. A customer receives the capability to retain and operate the released artefact.

The software stack is globally sourced

The Model Factory is European-controlled integration, not an all-European software stack.

The technical report names PyTorch 2.14, DeepEP v2 and FlashAttention 4 in pre-training; the trainer builds on Aleph Alpha’s fork of TorchTitan; vLLM is used for rollout inference and public serving; and NVIDIA tooling is used for diagnostics around the production cluster.61

Savanna itself uses Flyte, Kubernetes, Kueue, Grafana and other components in its orchestration environment.6263

Most of these components are open-source and can in principle be retained and modified. That reduces legal dependence on a single software vendor, but it does not make migration cost negligible. High-performance attention kernels, MoE communication, quantization, profiling and distributed execution are tightly coupled to the accelerator platform used in production.

The correct description is therefore European ownership of the integration and training recipe over a globally developed software substrate.

The hard sovereignty boundary remains the accelerator stack

Kolibri’s production pre-training used 768 NVIDIA B200 GPUs; the post-training infrastructure also uses NVIDIA B300 hardware for the main RL setup.64 This is the clearest non-European dependency in the demonstrated production chain.

The issue is larger than the GPU brand. A modern AI training platform depends on accelerator silicon, HBM, advanced packaging, high-speed intra-node links, inter-node networking, drivers, compilers, communication libraries, optimized kernels, firmware, servers and datacentre operations. Aleph Alpha has demonstrated that a European engineering team can operate this stack effectively. It has not demonstrated that the EU can replace it with a European-designed and European-manufactured equivalent at comparable performance and scale.

This is why training location should not be equated with hardware sovereignty. Aleph Alpha states that teams in Germany developed Kolibri and that the model was trained on infrastructure in Germany and Finland under European and German law.65 The public Kolibri material reviewed here does not identify the legal owner of every physical B200 asset used in those facilities. What is strongly evidenced is Aleph Alpha’s operational control of the workload and training process; physical asset ownership is less completely disclosed.

External teacher models create a second upstream dependency

The hardware boundary is not the only foreign dependency. The SFT pipeline uses GLM-5.2, GLM-5.3 and Qwen3.8-27B as principal generators or regenerators of fine-tuning material.66 The broader pre-training pipeline also uses external model families for rephrasing, translation and judging tasks.

This dependency has a different failure mode from an API dependency at runtime. Losing access to an external teacher would not by itself remove information already incorporated into an existing checkpoint or synthetic datasets that Aleph Alpha is entitled to retain. The dependency nevertheless matters for reproductive sovereignty: the public evidence does not establish that Aleph Alpha could regenerate a functionally equivalent synthetic corpus if all non-European teacher models became unavailable.

The technical report’s transparency on this point is important because it prevents a misleading interpretation of we own the entire pipeline. The evidence supports ownership of the orchestration, curation and training process. It does not support exclusive European provenance of every model that contributed information to that process.

Corporate control is also becoming transatlantic

A further boundary concerns the organization that holds the engineers, Model Factory, data processes and model-development IP.

On 16 September 2026, Aleph Alpha and Cohere announced a definitive business-combination agreement. The planned combined company will operate globally as Cohere, with headquarters, R&D centres and leadership roles in Germany and Canada. The transaction remained subject to final regulatory approvals in the latest official announcement reviewed for this article.67

The companies say the structure contains safeguards and oversight mechanisms intended to preserve operational control and sovereignty requirements in both jurisdictions. The public announcement does not disclose enough detail to determine future ownership and control of specific Kolibri-related IP, repository access, training data or Model Factory decision rights.

The released weights are less sensitive to that uncertainty: a customer that already holds the checkpoint can continue to run that version. Future model-production capability follows the organization, engineers, compute allocations and IP that produce the next generation.

A layer-by-layer assessment

The evidence supports a strong but bounded sovereignty assessment.

Layer Evidence from Kolibri Remaining external dependency Assessment
Human capability German teams built two model generations and operate the end-to-end pipeline retention of team and future governance strong
Data engineering internal web pipeline, filtering, deduplication, PII redaction, mix search global source corpus strong process control
Synthetic-data production Aleph Alpha controls prompts, curation, audits and mixing GLM, Qwen and other external teacher models strong orchestration, incomplete provenance independence
Tokenizer internally developed UniBPE research ideas are global strong
Architecture internally selected through ablations research ecosystem is global strong
Training recipe proxy-model search, multi-stage curriculum, optimizer and merge strategy documented external software substrate strong
Model Factory Savanna integrates orchestration, CI, artefact lineage, evaluation and clusters internal code is not publicly transferred strong producer capability
Training systems engineering hundreds of B200s, profiling, scaling, automatic recovery NVIDIA hardware/software ecosystem strong operational capability
Physical compute workloads operated in Germany/Finland ownership details incomplete; accelerators non-European mixed
Accelerator/semiconductor technology no European substitute demonstrated NVIDIA and global semiconductor chain weak
Released model artefact downloadable FP8/BF16 weights none required for continued use of retained version strong
Customer inference self-hosting documented accelerator stack remains NVIDIA-centric strong provider independence, limited hardware independence
External reproducibility extensive report, model card and content summary no complete public Model Factory or training corpus partial
Corporate control German company at release pending Cohere combination future structure transatlantic
Table 4: Evidence-based sovereignty assessment for Kolibri as of 7 October 2026.

This table explains why both extreme readings are wrong:

  • Calling Kolibri not sovereign merely because it uses NVIDIA hardware ignores substantial European control over architecture, data transformation, training, evaluation and the final artefact.
  • Calling it fully sovereign ignores the accelerator stack, external teacher models, incomplete public reproducibility and the prospective change in corporate governance.

What the achievement means for European technology

Kolibri establishes several facts that matter beyond the model itself:

  1. Europe has at least one private organization that has demonstrated an integrated production pipeline for training a modern sparse foundation model from raw data through post-training and release. This is stronger evidence than the existence of a research checkpoint or a European-hosted foreign model because it includes architecture search, data-mix search, distributed training, failure recovery, post-training environments and repeated model generations.
  2. The Model Factory converts tacit research knowledge into a persistent engineering asset. Savanna’s versioned recipe, immutable artefacts, automated evaluations and lineage allow decisions to accumulate across model generations rather than being reconstructed for each run. Whether this capability remains in Europe is therefore strategically more important than the continued availability of any one checkpoint.
  3. Aleph Alpha demonstrates that deployment sovereignty can be transferred independently of full production openness. The customer can retain and self-host the weights even though the complete training factory is not open-sourced. This is a practical form of provider-exit resilience.
  4. The work makes the remaining gaps easier to identify. The missing pieces are not abstract. They are advanced accelerator and memory supply, a more hardware-portable high-performance software stack, upstream teacher-model substitutability, broader public reproducibility if that is a policy goal, and stable governance over the organizations that hold the model-production capability.

None of these conclusions requires predicting that Kolibri will become a global frontier model or that Aleph Alpha’s approach will dominate European AI. The demonstrated achievement is narrower and more durable: a European team has built and operated an industrial foundation-model production system at large scale and has documented enough of it to make the dependency boundaries visible.

What Kolibri does not prove

Kolibri’s significance becomes clearer when the sovereignty claim is bounded by what the project does not establish. The model demonstrates substantial European control over the intellectual and operational process that turns data, experiments and compute into a deployable foundation model; it does not demonstrate autonomy across every layer on which that process depends. The distinction is fundamental because the relevant object is the dependency graph behind the model, not the nationality of the finished checkpoint.

Most obviously, Kolibri does not demonstrate European semiconductor sovereignty. Its base-model training used 768 NVIDIA B200 accelerators, while the reinforcement-learning infrastructure uses B300s. The surrounding production stack depends on NVLink, InfiniBand and ConnectX networking as well as PyTorch, DeepEP, FlashAttention, vLLM and other software closely optimized for the NVIDIA execution environment.68 Open-source software reduces the possibility of a purely contractual veto because its code can be retained and modified, but it does not remove the engineering cost of migrating an optimized distributed-training system to a different accelerator architecture. Nothing in the Kolibri documentation establishes that the demonstrated training recipe could presently be reproduced at comparable scale, throughput and reliability on a European-designed and European-manufactured compute platform. The strongest external dependency therefore remains below the model-development layer, in accelerators, HBM, packaging, networking and semiconductor manufacturing.

Nor does Kolibri establish exclusively European provenance for its training information. Aleph Alpha controls the acquisition, filtering, deduplication, synthesis, mixing and validation pipelines, but much of the underlying corpus comes from external sources; in the documented 20T-token pre-training mix, external datasets account for 64.3% of token presentations. Synthetic data introduce another dependency. The technical report explicitly documents external models in the data-production chain, including Gemma for English synthetic pre-training data, Mistral-Nemo for German rephrasing, Qwen3-32B for quality annotation, and GLM-5.2, GLM-5.3 and Qwen3.8-27B as principal generators or regenerators of supervised fine-tuning data.69 This does not make Kolibri a derivative checkpoint of those models: Aleph Alpha trained its own base model and controls the transformation pipeline. It does mean, however, that control over the pipeline is not equivalent to exclusive control over the provenance of every informational input.

The open-weight release also stops short of transferring Aleph Alpha’s complete model-production capability to the public. Apache 2.0 applies to the released weights and configuration files, giving a customer substantial freedom to retain, deploy and adapt the resulting artifact. It does not make Savanna, the complete training corpus, experiment history, data-generation infrastructure or the full production implementation publicly reproducible. Aleph Alpha retains what might be called producer reproductive sovereignty, whereas a holder of the checkpoint primarily obtains artifact and deployment sovereignty. This distinction matters because the ability to operate Kolibri-1 is materially different from the ability to independently reproduce Kolibri-2.

The geographical location of the training infrastructure should likewise not be confused with ownership of the physical assets. Aleph Alpha states that the model was developed in Germany and trained on infrastructure in Germany and Finland, and its technical documentation provides strong evidence that the company controlled scheduling, training, checkpoint recovery, evaluation and lineage. The reviewed sources do not, however, identify the legal owner of every B200 cluster used in the training process or fully disclose the provider structure behind that capacity. European jurisdiction and operational control are therefore established more strongly than European ownership of the physical compute estate.

Transparency has a similar boundary. Kolibri is accompanied by an unusually detailed technical report, model card, training-content summary, architectural specification and benchmark record, and Aleph Alpha has publicly committed to European transparency frameworks. That does not imply that every model output already carries machine-verifiable provenance or can be reliably identified as AI-generated. The relevant model documentation explicitly treats output detectability as an unresolved downstream concern. Transparency about how a model was built and technical provenance of what a model subsequently generates are separate properties.

Corporate sovereignty is also no longer reducible to Aleph Alpha’s historical German identity. The announced business combination with Cohere would create a transatlantic organization with operations, leadership and R&D in Germany and Canada. The companies state that sovereignty safeguards will be maintained, but the public material does not yet expose enough of the resulting governance structure to determine control of Kolibri-related IP, repositories, training data, Model Factory access or veto rights over future technology transfer. The appropriate conclusion is therefore not that European control has disappeared, but that exclusive European control over future model generations is unresolved.

Finally, Kolibri does not prove that European sovereign models have already reached the global capability frontier. Aleph Alpha’s principal claim is more specific: within its evaluated set, Kolibri occupies a favorable Pareto frontier between serving cost and model quality, particularly given its 3.46B active parameters per token. Its own broader benchmark tables nevertheless include stronger dense models on the overall English and German aggregates. The model is therefore evidence of efficient and competitive European model engineering, not evidence that Kolibri has reached the global capability frontier; Aleph Alpha’s own comparison already contains a stronger dense model on the aggregate English and German evaluations.

Taken together, these limitations identify three different notions that should not be collapsed into the single adjective sovereign. Kolibri provides strong evidence of model-production sovereignty: a European organization controls architecture, data transformation, experimentation, training, post-training and evaluation. Its downloadable weights provide substantial deployment sovereignty, because customers can retain and operate the model without depending on Aleph Alpha’s inference API. What it does not yet establish is full-stack reproductive sovereignty: the demonstrated ability to build successor models while independently replacing critical external accelerator technology, semiconductor manufacturing, upstream teacher models, physical compute supply and organizational control.

Those qualifications do not weaken the technical achievement. They define it. Kolibri demonstrates that European control is strongest at the intellectual and operational layers of foundation-model production, falls sharply at the physical compute layer, and then increases again once the finished weights are transferred to the customer. The remaining sovereignty problem is therefore no longer whether Europe can build a serious foundation model. It is whether Europe can preserve and reproduce that capability when one of the external layers on which the next model generation depends becomes unavailable.

Conclusion

Kolibri is a serious European AI achievement because it makes visible a capability that is more important than a single model release. Aleph Alpha has documented a pipeline that begins with hundreds of trillions of raw-token candidates, reduces them through curation and deduplication, searches data mixtures with proxy models, trains a sparse 78.1B-parameter architecture across nearly 24T token presentations, adapts it to 256k native context, generates and filters hundreds of billions of SFT token positions, trains on more than 1.2 million internally curated RL tasks, evaluates checkpoints continuously, survives large-cluster failures automatically, and publishes a self-hostable final artefact.

The Model Factory behind that process is the most sovereignty-relevant part of the system. It turns model development into a versioned, testable, traceable production process and allows successive generations to inherit engineering knowledge rather than rebuild it manually.

Aleph Alpha’s transparency choice is substantial but deliberately incomplete. The weights and configuration are released under Apache 2.0, accompanied by a long technical report, detailed model card and training-content summary. The complete production code, model factory, training corpus and associated IP are not transferred under that license. Kolibri is therefore open-weight and highly documented, not a fully open-source reproduction package.

The same precision is needed for the sovereignty claim. Aleph Alpha controls most of the intellectual and operational model-development pipeline and gives customers strong deployment autonomy. But the demonstrated system still depends on NVIDIA accelerators, an international semiconductor and software ecosystem, external teacher models for some synthetic data, and a corporate structure subject to a pending transatlantic business-combination agreement.

The evidence therefore supports a bounded conclusion:

Kolibri demonstrates substantial European sovereignty over the production and deployment of a foundation model, but not sovereignty over the complete technological stack that makes that production possible.

For the future of European technology, that distinction is useful because it turns AI sovereignty from a political slogan into an engineering map. Europe can now point to a demonstrated model factory, a demonstrated engineering team, a demonstrated data pipeline and a deployable open-weight artefact. The unresolved work lies below and around them: compute hardware, semiconductor supply, upstream model dependencies, external reproducibility and long-term organizational control.

That is a more demanding standard than simply asking whether Europe has an LLM. It is also a more useful one.

Appendix: how far is Kolibri from the global model frontier?

The preceding analysis establishes Kolibri as a significant European model-production achievement. A different question is whether it is already frontier-competitive in raw model capability. Sovereignty and capability are independent variables: Europe may control a model that remains materially behind the best American or Chinese systems, while a technically superior foreign model may offer little strategic control to its European user.

This appendix therefore asks a narrower question:

As of 7 October 2026, how large is the capability distance between Kolibri and the global frontier?

There is no direct authoritative answer. Kolibri was released on 3 October 2026 and, as of the date of this appendix, it does not have a published Artificial Analysis Intelligence Index score or an Arena placement comparable with the newest U.S. and Chinese systems. A direct leaderboard comparison with Claude Opus 5.5, GPT-6 Astra, GLM-5.3 or Kimi K3 is consequently unavailable.

There is, however, enough overlap to estimate the capability tier in which Kolibri plausibly sits without pretending that an inferred score is a measured one. Aleph Alpha evaluated Kolibri against Qwen3.8 27B, Qwen3.6, Qwen3.5, Nemotron 3 Super, Mistral Small 4, GLM-4.7 Flash, GPT-OSS 120B and several other models. Many of those same releases have Artificial Analysis benchmark values, although some of those values are explicitly estimated by Artificial Analysis rather than produced by a completed independent run. They can therefore be used as bridge models between Aleph Alpha’s evaluation suite and the current independent frontier.707172

The result is reasonably clear, but it must be stated with the appropriate uncertainty. A linear calibration over the overlapping models places Kolibri at about 21.4 on the current Artificial Analysis scale; several pairwise interpolations suggest a broader 20–26 envelope. That range is a heuristic cross-benchmark estimate, not an Artificial Analysis result and not a statistical confidence interval. At the same time, the leading Chinese open-weight models score about 44–45, GPT-6 Astra (Max) scores 53, and Claude Opus 5.5 (Max) scores 58.73747576

The implied gap is large, but it is highly non-uniform. On long-context reasoning, mathematics and conventional code benchmarks, Kolibri is considerably closer to the strongest models tested in Aleph Alpha’s own harness. The difference becomes much larger on difficult long-horizon agentic tasks and on broad closed-book knowledge reliability.

The independent frontier in October 2026

Artificial Analysis is useful here because its current Intelligence Index is intentionally broader and more agentic than conventional academic benchmark bundles. Version 4.3.2 combines ten evaluations spanning professional work, workflow automation, terminal operation, scientific coding, Humanity’s Last Exam, document work, difficult reasoning, knowledge reliability and long-context reasoning.77

Table Table 5 gives the reference points used in this appendix. Kolibri’s range is the inference developed below; the other Intelligence Index values are current Artificial Analysis values as accessed on 7 October 2026.

Model Origin Availability Total / active parameters Context AA Intelligence Index Evidentiary status
Kolibri Germany / Aleph Alpha open weights, Apache 2.0 78.1B / 3.46B 256k trained; up to 1M evaluated ≈20–26 cross-benchmark inference; not independently scored
Mistral Medium 3.5 France / Mistral AI open weights 128B / not disclosed 256k 14 Artificial Analysis result
Mistral Large 4 Preview France / Mistral AI API preview; open-weight release announced for end of October 1.05T / 49B 1M documented; 524k in the evaluated API endpoint 38 Artificial Analysis result
GPT-OSS 120B United States / OpenAI open weights 117B / 5.1B 128k 12 Artificial Analysis result
Qwen3.8 27B (xhigh) China / Alibaba open weights, Apache 2.0 27B / 27B 256k 34 Artificial Analysis result
GLM-5.3 Max China / Z.ai open weights, GLM-5.3 License 753B / 40B 1M 45 Artificial Analysis result
Kimi K3 Max China / Moonshot AI open weights, Kimi K3 License 2.8T / 104B 1M 44 Artificial Analysis result
GPT-6 Astra (Max) United States / OpenAI proprietary undisclosed 1M 53 Artificial Analysis result
Claude Opus 5.5 (Max) United States / Anthropic proprietary undisclosed 1M 58 Artificial Analysis result
Table 5: Reference capability levels used to estimate Kolibri’s position. Artificial Analysis values are a time-stamped 7 October 2026 snapshot; Kolibri’s value is an inferred interval, not an Artificial Analysis score.

Artificial Analysis currently places GLM-5.3 Max at 45, Kimi K3 Max at 44 and Qwen3.8 27B in its xhigh configuration at 34. Mistral Large 4 Preview, released on 6 October, scores 38, materially changing the European comparison even though its public weights are not yet available.7879808182 At the proprietary frontier, GPT-6 Astra (Max) scores 53 and Claude Opus 5.5 (Max) scores 58.8384

These numerical distances must not be interpreted as percentage differences in intelligence. The Artificial Analysis Intelligence Index is a composite benchmark score, not a ratio scale: a model scoring 50 is not meaningfully “twice as intelligent” as one scoring 25.

Why Kolibri cannot simply be assigned a leaderboard score

Aleph Alpha’s evaluation suite and the Artificial Analysis Intelligence Index measure overlapping but non-identical capability distributions.

Aleph Alpha’s post-training aggregate combines knowledge, mathematics, code, instruction following, agentic tool use, retrieval, grounding, industry RAG and other evaluations. Its stated purpose is to characterize a bilingual model designed for enterprise and regulated-domain deployment.85

Artificial Analysis v4.3.2 places substantial weight on difficult agentic and economically relevant work through evaluations such as AutomationBench-AA, Terminal-Bench 4.0, AA-Briefcase and GDPval-AA, alongside Humanity’s Last Exam, SciCode, AA-Omniscience and AA-LCR.86

Therefore,

an Aleph Alpha aggregate of 75.5 cannot be converted mechanically into an Artificial Analysis score.

A cross-benchmark calibration is nevertheless possible because the two systems contain a useful set of common model releases.

A transitive calibration

Aleph Alpha’s Table 28 reports Kolibri and a set of external baselines under one evaluation protocol. The current Artificial Analysis model pages provide corresponding Intelligence Index values for eleven of those releases. Some Artificial Analysis values are marked by the platform as estimates; this is shown explicitly in Table Table 6.8788

Bridge model Aleph Alpha Overall EN AA Intelligence Index AA status
Qwen3-Next 80B-A3B Thinking 62.4 11 estimated
Mistral Small 4 63.1 11 measured
GLM-4.5 Air 106B-A12B 64.4 11 estimated
GLM-4.7 Flash 30B-A3B 64.7 15 estimated
Nemotron 3 Nano 30B-A3B 65.6 9 measured
Qwen3.6 35B-A3B 71.4 18 measured
Gemma 4 26B-A4B IT 71.9 17 estimated
GPT-OSS 120B 72.3 12 measured
Nemotron 3 Super 120B-A12B 73.0 13 measured
Qwen3.5 35B-A3B 74.7 19 estimated
Qwen3.8 27B 80.2 34 measured
Kolibri 75.5 not measured target of the inference
Table 6: Bridge models used in the cross-benchmark calibration. Aleph Alpha scores come from its post-training Table 28. Artificial Analysis status distinguishes completed evaluations from values that Artificial Analysis itself labels as estimated.

The mapping is visibly noisy. For example, Nemotron 3 Super scores 73.0 in Aleph Alpha’s aggregate but only 13 on the Artificial Analysis index, whereas Qwen3.5 scores 74.7 and 19. That is expected: the suites weight different capabilities, and model profiles differ substantially across reasoning, knowledge and agentic work.

A least-squares fit over the eleven bridge points gives

\widehat{I}_{AA} = 0.972\,S_{\mathrm{Aleph}} - 52.0,

where S_{\mathrm{Aleph}} is Aleph Alpha’s English overall score and \widehat{I}_{AA} is only a cross-benchmark estimate of the Artificial Analysis Intelligence Index.

For Kolibri,

S_{\mathrm{Aleph}} = 75.5,

which gives

\widehat{I}_{AA} \approx 21.4.

The Pearson correlation over the eleven bridge points is approximately

r \approx 0.80.

That is strong enough to suggest a broad capability tier, but not strong enough to treat the fitted value as a substitute for an independent benchmark run.

Pairwise interpolation provides a useful robustness check. Interpolating Kolibri between Qwen3.5 and Qwen3.8 gives a value close to 21; using Qwen3.6 and Qwen3.8 produces roughly 25–26; using Nemotron 3 Super and Qwen3.8 gives roughly 20. Those calculations motivate the broader interval used here:

The overlapping-model evidence places Kolibri plausibly in the low-to-mid twenties on the current Artificial Analysis scale. A 20–26 interval is a useful heuristic envelope, not an official score and not a statistical confidence interval.

The calibration should be read with five limitations in mind:

  1. The two benchmark suites are not equivalent and do not assign the same weights to the same capabilities.
  2. Apparently identical model names can be evaluated with different sampling, prompting, reasoning-effort and serving configurations.
  3. Several Artificial Analysis bridge values are estimates rather than completed measurements.
  4. Artificial Analysis periodically changes its index and benchmark composition, so the mapping is inherently time-dependent.
  5. Eleven bridge models are enough to expose a rough relationship but not enough to justify a precise latent “intelligence” scale.

The purpose of the regression is therefore modest: to test whether Kolibri belongs broadly in the teens, twenties, thirties or frontier-forties-and-above. It should not be used to claim that Kolibri’s unseen Artificial Analysis score is exactly 21.4.

Shared benchmarks provide a partial consistency check

The transitive calibration would be much weaker if the two evaluation systems produced radically different values on every nominally shared metric. They do not.

Qwen3.8 27B is particularly useful because Aleph Alpha evaluates it in the same table as Kolibri and Artificial Analysis evaluates the corresponding release independently. The numbers are close on three shared benchmark names:

Evaluation Kolibri, Aleph harness Qwen3.8, Aleph harness Qwen3.8, Artificial Analysis GLM-5.3 Max, AA GPT-6 Astra (Max), AA Claude Opus 5.5 (Max), AA
Humanity’s Last Exam 21.5 35.6 34 42 55 61
AA-Omniscience Index −32.8 −9.5 ≈−10 14 43 46
AA-LCR 68.3 81.3 82 80 81 85
Table 7: Shared or near-shared benchmark names used as a consistency check. The Aleph Alpha and Artificial Analysis values should not be treated as strict reproductions because model configuration, harness details and reasoning settings can differ.

Aleph Alpha’s Qwen3.8 values are strikingly close to the current Artificial Analysis values.8990 This does not prove that the composite-score regression is valid, nor that the two organizations used identical inference configurations. In particular, Aleph Alpha serves Qwen3.8 with its documented default generation configuration, whereas the 34-point Artificial Analysis result is associated with the platform’s xhigh reasoning configuration. The agreement is therefore best treated as a partial consistency check, not as validation of an identity between the two evaluation systems.

The frontier values in the same table show why the capability gap is non-uniform. GLM-5.3 Max scores 42 on Humanity’s Last Exam, 14 on AA-Omniscience and 80 on AA-LCR; GPT-6 Astra (Max) scores 55, 43 and 81; Claude Opus 5.5 (Max) scores 61, 46 and 85.919293

Long-context reasoning: a moderate rather than categorical gap

Kolibri scores 68.3 on Aleph Alpha’s AA-LCR run. Qwen3.8 scores 81.3 in the same Aleph Alpha harness and 82 in Artificial Analysis. The current frontier values are about 80 for GLM-5.3, 81 for GPT-6 Astra and 85 for Claude Opus 5.5.9495969798

The exact point differences across organizations should not be overinterpreted because the harnesses are not guaranteed to be identical. The qualitative conclusion is nevertheless stable: Kolibri is materially below the leading systems, but the separation on long-context reasoning is much smaller than its inferred general-purpose composite gap.

That is consistent with the design of the model. Long context is not merely a serving-time extension: Kolibri’s base model is explicitly trained through a long-context stage at 256k sequence length and is evaluated by Aleph Alpha at up to one million tokens.99

For document-heavy workloads in public administration, legal analysis or industrial documentation, this is therefore one of the areas in which Kolibri is relatively close to much larger systems.

Mathematics is also comparatively strong

Under Aleph Alpha’s common evaluation protocol, the direct Kolibri–Qwen3.8 comparisons are:

  • AIME 2026: 96.0 vs 97.7;
  • GPQA Diamond: 84.3 vs 89.2;
  • AIME 2025: 96.9 vs 97.9.100

Kolibri also scores above the other MoE baselines in Aleph Alpha’s English AIME comparisons despite some of them activating substantially more parameters per token.101

These results do not establish frontier equivalence: they cover selected structured reasoning benchmarks and do not capture the full capability distribution. They do show that the aggregate distance from the frontier is not reproduced uniformly on every mathematical task.

Conventional coding is strong; long-horizon agentic execution is not

The same distinction appears in software engineering. Under Aleph Alpha’s evaluation harness:

Benchmark Kolibri Qwen3.8 27B
LiveCodeBench v6 85.9 93.8
HumanEval+ 92.7 94.7
SWE-Bench Verified 66.4 72.6
Table 8: Kolibri versus Qwen3.8 under Aleph Alpha’s common protocol on selected coding benchmarks.

The gap on these conventional code and software-engineering benchmarks is material but not enormous. On TerminalBench 2.1, however, Kolibri scores 27.7 while Qwen3.8 scores 76.8 in the same Aleph Alpha table.102

That difference is qualitatively larger. Terminal-style benchmarks require a model to interact with a computational environment over multiple steps, maintain state, recover from errors and complete operational tasks. They are therefore much closer to the behavior expected of autonomous coding agents than HumanEval-style function completion.

The evidence consequently suggests that one of Kolibri’s largest current capability deficits is long-horizon agentic execution, not basic code syntax or short-horizon algorithmic generation.

Broad knowledge reliability shows a much larger gap

Another major difference appears in AA-Omniscience. Artificial Analysis defines the Omniscience Index so that correct answers are rewarded, incorrect answers are penalized and abstention is not penalized; the metric can range from −100 to +100.103

Aleph Alpha reports Kolibri at −32.8 on the public-set index and Qwen3.8 at −9.5 under its own harness. Artificial Analysis reports Qwen3.8 at approximately −10, GLM-5.3 at +14, GPT-6 Astra at +43 and Claude Opus 5.5 at +46.104105106107108

Again, the Aleph and Artificial Analysis values should not be subtracted as if they came from one perfectly identical test environment. The direction and magnitude are nonetheless hard to ignore: Kolibri’s closed-book knowledge reliability and calibration are much weaker than those of the current proprietary frontier.

This distinction matters because Kolibri is designed for regulated enterprise and public-sector use. The model performs much better when answering from controlled retrieved context than its closed-book Omniscience result would suggest. Those are different capabilities.

A deployment architecture based on Kolibri can therefore rationally rely more heavily on retrieval, controlled knowledge stores and grounding rather than treating the model’s parametric memory as frontier-equivalent.

The Chinese frontier is the strategically harder comparison

If the only comparison were between Kolibri and proprietary systems such as Claude Opus 5.5 or GPT-6 Astra, the capability gap could be framed largely as a trade-off: Europe gains deployment sovereignty and accepts some loss of raw frontier capability.

The Chinese comparison is more demanding because GLM-5.3 and Kimi K3 are themselves downloadable open-weight systems. Their licenses are not identical to Apache 2.0—GLM-5.3 uses its own license and Kimi K3 has commercial conditions for some large-scale uses—but their model weights can be operated outside the provider’s hosted API.109110

Artificial Analysis currently places GLM-5.3 Max at 45 and Kimi K3 Max at 44. Mistral Large 4 Preview now scores 38, narrowing the measured European gap substantially, although its announced open weights had not yet been released as of 7 October 2026.111112113114

The strategic comparison is therefore not simply

European open models versus American closed APIs.

It is also

European open-weight models versus substantially stronger Chinese open-weight models.

This prevents the capability gap from being explained merely as the cost of choosing openness or local deployment. China demonstrates that downloadable weights and substantially higher general-purpose capability can coexist.

The model-scale difference is enormous

Part of the comparison is simply one of model scale. Kolibri has 78.1B total parameters and activates 3.46B per token.115 GLM-5.3 has 753B total parameters and 40B active; Kimi K3 has 2.8T total parameters and 104B active.116117

Comparison Total-parameter ratio vs Kolibri Active-parameter ratio vs Kolibri
GLM-5.3 / Kolibri ≈9.6× ≈11.6×
Kimi K3 / Kolibri ≈35.9× ≈30.1×
Table 9: Model scale relative to Kolibri. Parameter ratios are descriptive and do not imply proportional capability.

These ratios do not imply corresponding ratios in capability. Architecture, routing sparsity, data, training compute, post-training and inference-time reasoning all matter. They do, however, put Aleph Alpha’s result in perspective. Kolibri is being compared with Chinese systems that activate roughly 12× to 30× as many parameters per token. Its ability to approach Qwen3.8 on some mathematics, coding and long-context tasks is technically significant, but the larger deficit on broad agentic capability is not surprising.

Translating the inferred score into frontier distance

Using the deliberately broad 20–26 cross-benchmark envelope gives the following approximate distances:

Frontier reference AA Index Approximate distance above Kolibri
Qwen3.8 27B (xhigh) 34 8–14 points
Kimi K3 Max 44 18–24 points
GLM-5.3 Max 45 19–25 points
GPT-6 Astra (Max) 53 27–33 points
Claude Opus 5.5 (Max) 58 32–38 points
Table 10: Approximate frontier distance if Kolibri lies in the inferred 20–26 Intelligence Index envelope. Differences are benchmark points, not percentage deficits in intelligence.

The broad conclusion is robust to the uncertainty in the calibration. Even at the upper end of the envelope, the Chinese open-weight frontier would remain roughly twenty Intelligence Index points higher, while the strongest U.S. proprietary systems would remain more than thirty points higher.

The Mistral Large 4 release changes the European comparison. Kolibri’s inferred 20–26 range remains above earlier Mistral releases such as Medium 3.5, Small 4 and Large 3, but it is clearly below the independently measured 38 of Mistral Large 4 Preview.118119 Kolibri’s position remains an inference rather than an independent leaderboard result.

Arena provides an independent qualitative cross-check

Artificial Analysis is benchmark-driven. Arena provides a different signal based on human preferences. Kolibri is not yet present there either, so Arena cannot be used to assign it a capability score. It can, however, test whether the broader U.S.–China ordering is peculiar to Artificial Analysis.

The current evidence points in the same general direction. In Arena’s 2 October 2026 overall text leaderboard, Google’s Gemini 4 Argon High ranked first, Claude Opus 5.5 High ranked fourth and Kimi K3 Max ranked sixteenth among more than 400 models.120

In Arena’s 11 September open-weight snapshot, GLM-5.3 Max and Kimi K3 Max occupied the first two positions among open-weight models.121 Arena therefore provides a qualitative cross-check for one narrower proposition:

Chinese open-weight models are already competitive much closer to the top of the global model ecosystem than the currently measured European open-weight models.

That statement does not depend on converting Arena scores into Artificial Analysis scores, and Arena should not be treated as measuring the same latent quantity.

Kolibri is optimized for a different point on the Pareto surface

None of the frontier evidence makes Aleph Alpha’s Pareto argument false. It clarifies the objective function.

Kolibri is not designed to maximize aggregate benchmark performance regardless of serving footprint. Aleph Alpha reports that it rejected a substantially larger 123B candidate because the deployment penalty was disproportionate: on two H100 GPUs, the larger model could fit only three concurrent 256k-token requests, whereas the selected 78B design could accommodate eighteen and decode 28% faster.122

Kolibri therefore occupies an engineering point involving at least:

  • model capability;
  • active inference cost;
  • long-context operation;
  • German capability;
  • self-hostability;
  • grounded enterprise workloads; and
  • on-premises deployment.

GLM-5.3 and Kimi K3 operate at dramatically larger active model scales. Claude Opus 5.5 and GPT-6 Astra expose no downloadable weights.

The appropriate conclusion is therefore not that Kolibri simply “loses” because its aggregate capability is lower. It is that Kolibri’s particular combination of small active footprint, self-hostability and sovereignty currently carries a measurable general-capability opportunity cost, especially for difficult autonomous and agentic work. Mistral Large 4 shows that this trade-off should not be generalized to European model development as a whole.

The capability gap is not uniform

The evidence is more useful when decomposed by capability class.

Capability Kolibri relative to global frontier Evidence
German language specialized strength purpose-built tokenizer and >20% German pre-training share
Long-context reasoning moderate gap Aleph AA-LCR 68.3; frontier AA values around 80–85
Mathematical reasoning small-to-moderate gap on tested benchmarks AIME close to Qwen3.8 in Aleph’s common harness
Conventional code generation moderate gap strong LiveCodeBench and HumanEval+ results
Grounded enterprise retrieval specialized strength in Aleph Alpha tests strong customer-proxy and retrieval evaluations
Instruction following competitive with mid-tier open models in Aleph’s suite IFBench 78.1
Broad knowledge reliability large gap AA-Omniscience −32.8 in Aleph harness vs +14 China frontier and +43/+46 U.S. frontier on AA
Long-horizon agentic coding very large gap TerminalBench 2.1: 27.7 vs Qwen3.8 76.8 in Aleph’s common harness
General frontier composite large inferred gap inferred ≈20–26 vs 44–45 Chinese open-weight and 53–58 U.S. proprietary frontier
Table 11: Capability distance is highly task-dependent. Cross-organization comparisons are marked as such and should not be read as perfectly controlled head-to-head experiments.

Europe is therefore not “behind” in one uniform dimension. Kolibri demonstrates competitive engineering in efficient sparse architectures, German specialization, long context, local deployment and selected forms of mathematical and coding reasoning.

The evidence is also consistent with a substantially larger deficit in long-horizon agentic execution and broad closed-book knowledge reliability. The public record does not allow that difference to be attributed to one cause. Model scale, data, training compute, RL environments, post-training methodology and inference-time reasoning are all plausible contributors.

The gap exists above the semiconductor layer as well

This matters for the sovereignty thesis of the main article.

The analysis of Kolibri already identifies NVIDIA B200/B300 dependence as a major physical-layer limitation. The frontier comparison shows something different: compute sovereignty alone would not automatically erase the observed model-capability gap.

GLM-5.3 and Kimi K3 do not merely sit on different accelerator estates. Their released models operate at much larger total and active parameter scales and achieve substantially stronger results on several current agentic and knowledge benchmarks.

Closing the European gap therefore requires more than semiconductor policy. It plausibly requires, at minimum:

  1. scaling European base models while preserving deployment efficiency;
  2. larger and more sophisticated RL and agentic-environment programmes;
  3. stronger software-engineering and autonomous-tool-use training;
  4. better broad-knowledge reliability and calibration; and
  5. repeated, expensive model generations rather than a single successful checkpoint.

These are engineering implications of the observed capability profile, not direct measurements of the hidden training programmes of frontier vendors.

Kolibri’s Model Factory matters precisely because model capability is iterative. Its strategic value is not that Kolibri has already matched Claude, GPT, GLM or Kimi; it is that Aleph Alpha has demonstrated a version-controlled production system capable of running repeated model-development cycles.

A particularly informative comparison: Kolibri versus Qwen3.8

Qwen3.8 27B is the most useful bridge model in the comparison. It is a dense 27B model, so it activates almost eight times Kolibri’s 3.46B parameters per token. Aleph Alpha evaluates it directly, Artificial Analysis evaluates it independently, and the two organizations obtain very similar values on several overlapping benchmark names despite differences in inference configuration.

Qwen3.8 is also not the current Chinese frontier. Artificial Analysis places the 27B xhigh configuration at 34, while the much larger Qwen3.8 2.4T-A95B release reaches about 40; GLM-5.3 and Kimi K3 reach 45 and 44.123124125126

Figure 8: Artificial Analysis Intelligence Index for selected open-weight and frontier models. Source: Artificial Analysis, retrieved 7 October 2026.

The current capability ladder can therefore be represented approximately as follows:

Capability tier Representative models Interpretation
European sovereign baseline Kolibri Strong European model-production capability, but a substantial inferred gap remains on broad general and agentic capability
Intermediate Chinese open-weight tier Qwen3.8 27B Materially stronger general capability while remaining relatively compact and openly deployable
Chinese open-weight frontier Qwen3.8 2.4T-A95B, GLM-5.3, Kimi K3 Current leading open-weight capability tier, around 40–45 on the Artificial Analysis Intelligence Index
Proprietary U.S. frontier GPT-6 Astra, Claude Opus 5.5 Highest general capability in the comparison, but without open-weight deployment autonomy
Table 12: Approximate capability tiers relevant to the Kolibri comparison. The ordering represents broad aggregate capability rather than a strict ranking on every task.

The ordering in Table 12 is not strict for every task; it represents broad capability tiers visible in the current independent evaluations.

A nearer technical objective for Europe would therefore be a sovereign model that independently matches or exceeds Qwen3.8 27B across modern agentic and knowledge evaluations while retaining Kolibri’s German specialization and efficient deployment footprint. A subsequent milestone would be parity with the 40–45-point Chinese open-weight class. Only after those steps would the proprietary U.S. frontier become the immediate capability comparison.

Open weights change the strategic meaning of the frontier

The strongest U.S. models remain proprietary, and their total parameter counts and training compute are not publicly disclosed. That prevents a meaningful parameter-efficiency comparison with Kolibri.

The Chinese systems are different. Their downloadable weights make the deployment side of the comparison inspectable and allow third-party operation outside the originating provider’s API, subject to their respective license terms.127128

This creates an uncomfortable but important sovereignty observation. Europe has strong legal and political incentives to reduce dependence on foreign AI infrastructure, yet the highest-capability openly deployable models in the present comparison are Chinese rather than European.

European technological sovereignty therefore cannot be achieved merely by choosing open weights instead of closed U.S. APIs. Without a sufficiently capable European open-weight alternative, an organization seeking both high capability and local deployment may simply move part of its dependency from an American service provider to a Chinese model artifact.

That can increase runtime autonomy. It is not equivalent to European technological sovereignty.

What would count as closing the gap?

A credible next European milestone would not require beating every frontier model on every public benchmark. It would require eliminating the obvious capability-tier separation.

For a successor to Kolibri, evidence of convergence would include:

  • an independent Artificial Analysis or equivalent result around the 40+ tier, rather than an inferred low-twenties position;
  • agentic terminal and software-engineering performance comparable with the leading Chinese open-weight models;
  • AA-Omniscience moving from strongly negative into positive territory under a controlled independent evaluation;
  • preservation of Kolibri’s German, long-context, grounding and self-hosting strengths; and
  • substantially better deployment efficiency than the much larger 40–104B-active Chinese systems.

That would be strategically important even if the best proprietary U.S. models remained ahead. It would mean that Europe possessed an open-weight model near the current open frontier while retaining European control over a substantial part of the production pipeline.

Kolibri does not yet demonstrate that. It demonstrates something more foundational: that Aleph Alpha has built much of the machinery required to attempt repeated model generations at this level.

Conclusion: a real European model factory, but not yet a frontier-equivalent model

The most defensible conclusion is more demanding than either the launch narrative or a dismissive comparison based on one leaderboard.

Kolibri is technically impressive for its scale. It activates only 3.46B parameters per token yet reaches 96.0 on AIME 2026, 84.3 on GPQA Diamond, 85.9 on LiveCodeBench v6 and 68.3 on Aleph Alpha’s AA-LCR evaluation, while performing strongly on several industrial retrieval workloads.129

Its small active footprint, German specialization, long context, downloadable weights and enterprise-oriented design make it materially different from a generic leaderboard model.

But efficiency should not obscure capability distance. Using the eleven models evaluated both by Aleph Alpha and represented on Artificial Analysis as a transitive calibration, Kolibri most plausibly lies in approximately the 20–26 capability tier of the current Artificial Analysis Intelligence Index. The point estimate from a simple linear fit is about 21.4, with a bridge-model correlation of about 0.80. Those figures are methodological aids, not independent measurements.

The leading Chinese open-weight systems are currently around 44–45. GPT-6 Astra (Max) is around 53. Claude Opus 5.5 (Max) reaches 58.130131132133

The direct and near-direct benchmark evidence helps explain where the gap comes from. Kolibri is relatively close on some long-context and mathematical evaluations and remains competitive on conventional coding benchmarks. It falls much further behind on broad knowledge reliability and difficult long-horizon agentic execution.

This refines the sovereignty judgment of the main article:

Europe has demonstrated both a serious sovereign model-production stack through Kolibri and, with Mistral Large 4 Preview, a European model much closer to the global capability frontier. It has not yet demonstrated a publicly deployable European open-weight model at parity with the leading Chinese open-weight systems or the strongest U.S. proprietary models.

Kolibri therefore answers the question Can Europe build its own modern foundation model? increasingly convincingly.

The next question is harder:

Can Europe scale that sovereign model-production capability until choosing European technological control no longer requires accepting a material capability gap?

As of 7 October 2026, the evidence says not yet. The gap is nevertheless measurable across several overlapping evaluations and, more importantly, decomposable into specific technical capability classes.

Appendix: European model sovereignty after Mistral Large 4, and Kolibri’s upstream training supply chain

The analysis in How Sovereign Is Kolibri? treats sovereignty as a layered property rather than a label attached to the nationality of a developer or the location of a datacentre. This appendix extends that analysis in two directions.

First, the European comparison changed materially on 6 October 2026, when Mistral AI launched the public preview of Mistral Large 4, internally nicknamed Le Chonk. Mistral Large 4 is a 1.05-trillion-parameter, natively multimodal Mixture-of-Experts model with 49 billion active parameters per token, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned datacentres in Europe. Artificial Analysis currently scores the preview at 38 on its Intelligence Index.134135136

Second, Kolibri exposes a different sovereignty boundary: its training-information supply chain is not exclusively European. Aleph Alpha’s technical report documents GLM-5.2, GLM-5.3 and Qwen3.8-27B as principal generators or regenerators of supervised fine-tuning data, with other external models used for pre-training-data generation, scoring and dataset construction.137 The September 2026 NSA/CISA/FBI advisory AA26-251A makes that provenance question more consequential, but it does not establish that Kolibri was compromised, poisoned or trained on unlawfully extracted model outputs.138

These two developments point to the same analytical requirement: sovereign AI must be assessed separately across model-development control, compute control, accelerator dependence, training-information provenance, public reproducibility, customer deployment autonomy and corporate governance.

European-developed foundation models: the comparison after Mistral Large 4

Several contemporary European projects meet a stronger criterion than European hosting or European fine-tuning: they have been trained from scratch by European organizations and therefore demonstrate indigenous model-development capability.

The most relevant comparators are summarized in Table 13.

Model or programme Developer / institutional base Main technical characteristics Training compute and location Openness and deployment Principal sovereignty limitation
Mistral Large 4 Preview Mistral AI, France 1.05T total / 49B active MoE; 52B active including embeddings/output layers; 1.6B vision encoder; natively multimodal; 1M documented context; 160+ languages trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres; preview API served on the same infrastructure Mistral announces an open-weight release by end of October; as of 7 Oct the public weights are not yet available and Artificial Analysis classifies the preview as proprietary NVIDIA/semiconductor dependence; public weight release and license still pending; complete architecture/post-training details not yet published
Kolibri Aleph Alpha, Germany 78.1B total / 3.46B active MoE; German and English; 20T pre-training tokens plus mid-training and long-context stages; 256k trained context, evaluated up to 1M 768 NVIDIA B200 GPUs for pre-training; development in Germany; training infrastructure in Germany and Finland FP8/BF16 weights already downloadable; self-hosting supported; detailed technical report and model card NVIDIA stack; external synthetic-data generators; incomplete public reproduction of Model Factory; prospective Cohere governance
Apertus ETH Zurich, EPFL, CSCS, Switzerland 8B and 70B multilingual models; 70B trained on approximately 15T tokens 4,096 NVIDIA GH200 accelerators on CSCS Alps unusually extensive release of weights, training code, reconstruction tooling, recipes and intermediate checkpoints Switzerland is European but outside the EU; accelerator technology remains non-European.139
EuroLLM-22B European research consortium 22B dense model; 35 languages including all 24 official EU languages; approximately 4T pre-training tokens approximately 400 NVIDIA H100 GPUs on MareNostrum 5 through EuroHPC open research-oriented release with supporting assets much smaller capability/scale than ML4; NVIDIA dependence
Salamandra / ALIA Barcelona Supercomputing Center, Spain 2B, 7B and 40B decoder-only models; 35 European languages; 12.875T-token pre-training corpus documented for the family European public research infrastructure Apache 2.0; weights, training scripts and configurations published public multilingual capability, but not frontier-scale industrial autonomy
Teuken-7B OpenGPT-X / Fraunhofer-led consortium, Germany 7B multilingual decoder model; all 24 official EU languages; current v0.6 base model reports 6T pre-training tokens European research programme downloadable model variants and detailed training documentation; licensing varies by release moderate scale; not a frontier-capability substitute
OpenEuroLLM EU-funded multi-institution programme programme rather than one completed frontier model; develops models, data, tooling and training methodology more than 10 million GPU-hours of strategic access across LUMI, Leonardo, JUPITER and MareNostrum 5 explicit objective of open models, data catalogues, tools, recipes and intermediate results programme outputs are not equivalent to an already completed frontier model; accelerators remain externally sourced
Table 13: European-developed models and programmes relevant to technological-sovereignty analysis as of 7 October 2026.

The comparison now extends from public multilingual and reproducible models to frontier-adjacent industrial capability at trillion-parameter scale. The relevant distinction is therefore no longer whether Europe can train foundation models at all, but which layers of the resulting capability each project controls.

Mistral Large 4 changes the European baseline

Mistral Large 4 is the most consequential update to the European comparison. Mistral’s official documentation describes ML4 as a granular Mixture-of-Experts model with approximately 1.05 trillion total parameters, 49 billion active parameters per token and a 1.6-billion-parameter vision encoder. It is natively multimodal and supports a documented context window of up to one million tokens.140

Mistral’s launch announcement provides an unusually strong infrastructure-sovereignty claim: the model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacentres in Europe, and the public-preview API is served on the same infrastructure. Mistral further states that it operates a European deployment end-to-end, independently of other digital service providers and under European law.141

That is materially stronger evidence of physical infrastructure control than was available for Mistral Large 3 and, on the public record, clearer than the legal-ownership picture around Kolibri’s B200 training infrastructure.

The capability shift is equally material. Artificial Analysis currently gives Mistral Large 4 Preview an Intelligence Index score of 38, compared with 14 for Mistral Medium 3.5, 11 for Mistral Small 4 and 9 for Mistral Large 3. Artificial Analysis describes ML4 as the most capable model from outside the United States and China in its current index.142

The same independent evaluation reports 60% on AutomationBench-AA, 27% on Terminal-Bench 4.0, 54% on SciCode, 35% on Humanity’s Last Exam, −5 on AA-Omniscience and 81% on AA-LCR v1.1.143 Mistral’s own announcement additionally reports a 49.8% Coding Agent Index, 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4.0 and strong results in cybersecurity, finance, legal work and multimodal grounding.144

These figures should be interpreted carefully. Some come from Mistral’s own evaluation programme, some from third parties named by Mistral, and some from Artificial Analysis. The most useful independent scalar for the comparison here is the Artificial Analysis score of 38.

The update materially narrows the measured European capability gap. Contemporary Chinese open-weight leaders such as GLM-5.3 Max and Kimi K3 Max sit around 45 and 44 on the same Artificial Analysis index, so ML4 is no longer separated from that class by the twenty-plus-point gap that characterized the earlier European comparison. It remains below those systems in the aggregate, but it now sits in the same broad frontier-adjacent band rather than a different capability tier.145

The important caveat: Mistral Large 4 is still a preview

The sovereignty consequences of ML4 should not be overstated before the public weight release. As of 7 October 2026, Mistral provides a public API preview. The company says the weights will be released by the end of October and has created an upcoming Hugging Face release page, but the public checkpoint is not yet generally downloadable.146147

That means two statements must be distinguished:

Mistral Large 4 is designed and announced as an open-weight model.

and

Mistral Large 4 weights are publicly downloadable today.

The first is supported. The second is not yet true as of this appendix’s date.

Artificial Analysis accordingly labels the current preview as proprietary and reports Open Source (Weights): No for the API-accessible checkpoint.148 That classification reflects present availability rather than contradicting Mistral’s announced release plan.

The exact public weight license is also not stated in the launch announcement reviewed here. Mistral says it will release the weights and further architecture/post-training details later in October; until that happens, the appropriate sovereignty assessment is announced deployment sovereignty, not yet fully exercisable public deployment sovereignty.

This distinction matters because Kolibri’s weights are already downloadable. Today, Kolibri therefore gives a customer stronger immediately exercisable provider-exit autonomy, even though ML4 is a much larger and more capable model.

Mistral’s Model Factory analogue is also becoming visible

The ML4 announcement provides evidence that Mistral’s capability is not only a single large pre-training run.

Mistral says the model uses the same training, customization and reinforcement-learning environment offered to customers through Mistral Forge.149 Its RL library exposes a composable environment interface spanning chat, scientific problem solving, safety alignment, factuality and long-horizon tool use. The environments share resources such as code sandboxes, web search and external APIs; verification can combine reward models, unit tests, LLM judges and static checks.

Mistral further describes asynchronous post-training in which an autoscaling fleet of actors generates tens of thousands of rollouts in parallel while training continues. The system supports very long trajectories with rollout budgets reaching millions of tokens across compactions. At the currently disclosed scale, a single RL run uses about 3,000 GPUs, generates roughly 33 billion rollout tokens per day, and retains approximately 16 billion trainable completion tokens per day after filtering and masking.150

This is not the same transparency level as Aleph Alpha’s Savanna documentation: Mistral has not yet published an equivalent end-to-end technical report for ML4, and says further architecture and post-training details are still forthcoming. But it is substantial evidence that Mistral controls an industrial model-production system rather than a one-off checkpoint.

European sovereignty is becoming differentiated

Mistral Large 4, Kolibri and Apertus should not be collapsed into a single sovereignty ranking because they demonstrate different strengths.

Mistral Large 4 currently provides the strongest evidence of European industrial scale, owned European compute infrastructure and independently measured frontier-adjacent capability. Kolibri provides unusually detailed evidence of the internal model-production pipeline and already transfers the released model artifact to customers through downloadable weights. Apertus goes furthest on public reproducibility, publishing not only weights but extensive reconstruction and training resources.

The significance of the European model ecosystem is therefore its emerging division of capabilities rather than the existence of one uniquely sovereign model. The broader comparison in Table 14 makes those differences explicit.

EuroLLM: multilingual capability as sovereignty

EuroLLM-22B addresses another strategic layer: language.

The project covers 35 languages, including all 24 official languages of the European Union, and was trained on approximately four trillion tokens using about 400 NVIDIA H100 GPUs on MareNostrum 5 through EuroHPC.151

This is not frontier scale in the sense of Mistral Large 4. Its strategic importance lies elsewhere.

A multilingual European model reduces the risk that linguistic capability is determined entirely by the data distributions, tokenizers and commercial priorities of non-European developers. A model that handles European languages natively can support public administration, education, legal work and regional enterprise without treating smaller European languages as incidental additions to an English-centric system.

The same point now applies at greater scale to Mistral Large 4: Mistral says a significant share of its training data is multilingual across more than 160 languages, including every official language of the European Union.152

EuroLLM therefore remains important not because it is Europe’s largest model, but because it represents a research programme whose design objective is explicitly European multilingual coverage and open research infrastructure.

Salamandra, ALIA and Teuken: public multilingual capacity below frontier scale

Salamandra/ALIA and Teuken should not be omitted simply because they are much smaller than ML4.

The Barcelona Supercomputing Center’s Salamandra family spans 2B, 7B and 40B parameters and was trained from scratch on 12.875 trillion tokens across 35 European languages and code. Its model card states that the family is released under Apache 2.0 and that training scripts and configuration files are public.153

Teuken, developed in the OpenGPT-X ecosystem led by Fraunhofer with Forschungszentrum Jülich, TU Dresden and DFKI, is a from-scratch multilingual model covering all 24 official EU languages. Current v0.6 model cards report 6T pre-training tokens for the 7B base model; earlier commercial variants were released under Apache 2.0, while later research releases use different licensing.154

Neither family is evidence of frontier-scale autonomy. They demonstrate something infrastructural: European organizations can collect multilingual data, train tokenizers and models from scratch, document their training process and distribute artifacts for third-party use.

That human and institutional capability remains a strategic asset even after much larger models become available.

OpenEuroLLM: sovereignty as a distributed capability

OpenEuroLLM differs from the other entries because it is better understood as a model-production programme than as a single finished flagship.

Its objective is to construct openly available European foundation models together with the datasets, tools, recipes, catalogues and intermediate results required to reproduce and extend them.155

The programme has received more than ten million GPU-hours of strategic EuroHPC access across LUMI, Leonardo, JUPITER and MareNostrum 5.156

That changes the locus of sovereignty. Rather than concentrating the complete production capability inside a single private firm, OpenEuroLLM attempts to distribute model-development knowledge across universities, supercomputing centres, research institutes and companies.

Mistral Large 4 makes this model no less relevant. A private European champion can create frontier-scale capability; a distributed public programme addresses a different failure mode—the concentration of tacit model-building knowledge inside one corporate organization.

Mistral Large 4 strengthens infrastructure sovereignty, but not semiconductor sovereignty

The European projects differ in scale, governance and openness, but they converge at the accelerator layer:

  • Kolibri uses NVIDIA B200 and B300 systems.
  • Mistral Large 4 was trained on 3,800 NVIDIA Grace Blackwell GPUs.157
  • Apertus uses NVIDIA GH200 systems.
  • EuroLLM-22B uses NVIDIA H100s.
  • EuroHPC systems combine European public infrastructure policy with accelerator technology largely designed and supplied outside Europe.

Mistral nevertheless closes an important part of the infrastructure-control gap. Its launch material states that ML4 was trained in Mistral-owned European datacentres and that the European service is operated end-to-end without another digital service provider.158 That gives Mistral direct control over datacentre operations, compute allocation and model execution while leaving the underlying Grace Blackwell accelerators, HBM, packaging and semiconductor supply chain externally dependent.

The appropriate description is therefore European infrastructure sovereignty built on a non-European accelerator substrate: materially stronger than European hosting on a foreign hyperscale cloud, but not full-stack technological autonomy.

Updated sovereignty comparison

The comparison can therefore be expressed descriptively rather than as a synthetic score.

Model / programme European model-development control Public reproducibility European infrastructure control Accelerator sovereignty Customer deployment sovereignty Governance
Mistral Large 4 Preview strong, frontier-scale partial and currently incomplete; detailed architecture/post-training disclosure pending strong: Mistral-owned European datacentres stated explicitly low announced strong, but public weights not yet released on 7 Oct French private company
Kolibri strong partial: extensive report, incomplete public Model Factory substantial operational control; physical ownership partly undisclosed low strong now through downloadable weights German company at release; announced Cohere combination creates transatlantic future governance
Apertus strong unusually strong strong public Swiss compute control low strong public academic / supercomputing institutions
EuroLLM-22B strong at 22B scale comparatively strong research openness strong EuroHPC participation low strong European research consortium
Salamandra / ALIA strong at moderate scale strong public research orientation low / externally dependent strong Spanish public research institution
Teuken strong at moderate scale strong for released research artifacts European research ecosystem low / externally dependent strong for downloadable releases German public/private research consortium
OpenEuroLLM programme designed for distributed reproductive capability central project objective strong EuroHPC allocation currently low intended to be strong EU-funded multi-institution consortium
Table 14: Sovereignty mechanisms differ even among genuinely European-developed models.

Mistral Large 4 changes one part of the earlier conclusion decisively: Europe now has a private model developer that combines a large proprietary compute estate in Europe with a model that independent evaluation places much closer to the global frontier.

What it does not change is the accelerator boundary.

Nor does it eliminate the separate question exposed by Kolibri: the provenance of the models used inside the data-production pipeline.

Kolibri’s synthetic-data supply chain

Aleph Alpha’s technical report is unusually explicit about external teacher models. Section 3.1.1 states that the fine-tuning pool combines open datasets with synthetic data and that some open datasets have their completions regenerated after failing Aleph Alpha’s audits.

It then identifies the main models used for those tasks:

The main models we use to generate this data, and to regenerate parts of the open datasets, are GLM-5.2, GLM-5.3 and Qwen3.8-27B.

Aleph Alpha adds that different generators are chosen for different tasks and explicitly observes that they have different values and biases, which the company says it consciously selects or avoids.159 This is a consequential disclosure. Kolibri is not a fine-tune of GLM or Qwen. Its base model is Aleph Alpha’s own and was trained from random initialization. The correct relationship is instead:

Kolibri is fine-tuned on a mixture that contains synthetic examples generated or regenerated by GLM and Qwen models.

External models also appear elsewhere in the data pipeline.

External model Documented function in Kolibri production Stage
GLM-5.2 principal SFT generation/regeneration; task-specific generation supervised fine-tuning / dataset construction
GLM-5.3 principal SFT generation/regeneration; reasoning, German and agentic generation supervised fine-tuning / dataset construction
Qwen3.8-27B principal SFT generation/regeneration; science, German chat and long-context generation supervised fine-tuning / dataset construction
Qwen3-32B LLM-as-judge annotations used to train text-quality classifiers pre-training-data curation
Gemma-4-26B-A4B rephrasing of deduplicated English Common Crawl material synthetic pre-training data
Mistral-Nemo-Instruct-2407 rewriting of organic German documents synthetic German pre-training data
Gemma-family judge judging generated examples in parts of the retrieval pipeline validation / dataset construction
Kimi-family model arbitration or judging in a long-context data-construction path validation / dataset construction
Table 15: External foundation models documented as upstream data-production or evaluation tools in Kolibri’s pipeline.

The dependency is therefore not incidental. External models participate as generators, regenerators, classifiers, judges and arbiters.

%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD

    R["Raw / open documents"]

    subgraph PRE["Pre-training data production"]
        G["Gemma-4-26B-A4B"]
        M["Mistral-Nemo"]
        QJ["Qwen3-32B"]
        P1["English rephrases"]
        P2["German rephrases"]
        P3["Quality labels / distilled classifiers"]

        G --> P1
        M --> P2
        QJ --> P3
    end

    subgraph SFT["SFT data production"]
        GLM52["GLM-5.2"]
        GLM53["GLM-5.3"]
        Q38["Qwen3.8-27B"]
        J["Other task-specific judges / arbiters"]
        SG["Synthetic / regenerated SFT examples"]
        JV["Judging / validation"]

        GLM52 --> SG
        GLM53 --> SG
        Q38 --> SG
        J --> JV
    end

    C["Aleph Alpha curation<br/>filtering · auditing · decontamination<br/>validation · mixture construction"]

    BT["Kolibri base training"]
    PT["Kolibri SFT + RL"]
    W["Released Kolibri weights"]

    R --> C
    P1 --> C
    P2 --> C
    P3 --> C
    SG --> C
    JV --> C

    C --> BT
    BT --> PT
    C --> PT
    PT --> W
Figure 9: Kolibri’s documented external-model supply chain. Pre-training synthesis and quality scoring are distinct from SFT generation and judging; Aleph Alpha controls curation, validation, mixing and downstream training.

Figure Figure 9 makes the sovereignty boundary visible without collapsing pre-training synthesis and SFT generation into one path.

Aleph Alpha controls the downstream selection, filtering, dataset composition, training and resulting model weights. It does not control the complete development history of every upstream model whose outputs enter that process.

Synthetic data are a material part of post-training

The scale of Kolibri’s synthetic-data programme makes the upstream-teacher question more than a marginal provenance issue. Aleph Alpha’s release announcement states that it generated approximately 174 billion synthetic tokens for supervised fine-tuning and, after filtering and combining them with permissively licensed open datasets, constructed an approximately 268-billion-token SFT mixture.160 The technical report separately summarizes the SFT programme as approximately 537 billion training-token presentations.161

These figures describe different stages of the pipeline rather than competing estimates of the same quantity. The 174B figure refers to generated synthetic material; the approximately 268B figure refers to one full packed SFT mixture/run; and the approximately 537B figure corresponds to the two full-length candidate SFT runs whose weights are combined in the final SFT checkpoint.

What the public documentation does not disclose is the contribution of each individual external teacher. There is no published breakdown showing what fraction of the generated material came from GLM-5.2, GLM-5.3 or Qwen3.8-27B. For capability analysis this may be secondary, but for provenance assurance it is significant: without that decomposition, the influence of a particular upstream model cannot be quantified directly from the public record.

Aleph Alpha does not blindly ingest teacher outputs

The use of external teacher models does not imply that their outputs are accepted uncritically. Aleph Alpha’s technical report describes a heterogeneous quality-control pipeline that includes dataset audits, regeneration of failed completions, language validation, benchmark decontamination in later-stage training data, task-specific verification, execution-based testing for software tasks, separate judges in some pipelines, reference-answer checks where available and extensive downstream model evaluation.162 The company also reports identifying incorrect science answers and datasets that degraded downstream performance, then regenerating those completions with alternative models. This is concrete evidence that the Model Factory can detect and correct at least some failures introduced during synthetic-data generation.

The assurance regime is not uniform across all datasets, however. In one German chat-generation path, GLM-5.3 generates reasoning answers while Qwen3.8-27B produces non-reasoning answers. Aleph Alpha notes that these samples lack reference answers and that it does not apply an additional LLM-quality filter in that path because experiments did not show sufficient benefit relative to the computational cost.163

The accurate characterization therefore lies between two extremes. Kolibri’s synthetic data are neither ingested without validation nor independently verified example by example. Aleph Alpha applies a risk- and task-dependent validation regime, with stronger assurance mechanisms where deterministic checks, references or additional judges are available and weaker guarantees where validation itself would require another probabilistic model.

CISA AA26-251A changes the provenance question

On 8 September 2026, the NSA, CISA and FBI published joint advisory AA26-251A, China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies.164 The advisory names six China-based AI companies, including Alibaba, Z.AI and Moonshot AI, and alleges systematic large-scale extraction of proprietary functionality or capabilities from U.S. frontier-model providers.

The overlap with Kolibri’s upstream model supply chain is organizational, not checkpoint-specific. Alibaba is associated with Qwen, and Kolibri uses Qwen3.8-27B as a principal SFT generator and Qwen3-32B as a pre-training-data annotation model. Z.AI develops GLM, and Kolibri uses GLM-5.2 and GLM-5.3 as principal SFT generators. Moonshot AI develops Kimi, and a Kimi-family model appears in Kolibri’s technical report in a narrower judging/arbitration role.

That overlap is relevant to provenance, but the evidentiary boundary is strict: AA26-251A does not establish that the exact GLM, Qwen or Kimi checkpoints used by Aleph Alpha contain material obtained through the campaigns described in the advisory, nor does it provide evidence that Aleph Alpha participated in those campaigns or that Kolibri is poisoned, backdoored or otherwise compromised.

Proposition Status
GLM-5.2, GLM-5.3 and Qwen3.8-27B are principal SFT generators/regenerators for Kolibri established by Aleph Alpha
Other external models generated, scored or judged Kolibri training material established
U.S. agencies allege industrial-scale unauthorized distillation by Z.AI and Alibaba established as a U.S. government allegation
The exact GLM-5.2/5.3 checkpoints used by Aleph Alpha contain the specific extracted material described by the advisory not established
The exact Qwen3.8-27B checkpoint used by Aleph Alpha contains the specific extracted material described by the advisory not established
Kolibri therefore contains identifiable Claude or GPT training examples not established
Aleph Alpha participated in the alleged extraction campaigns not established
Kolibri is poisoned, backdoored or compromised because it used these teachers no evidence in the reviewed sources
External-teacher provenance is relevant to a claim of fully sovereign model production supported as an architectural and supply-chain conclusion
Table 16: Evidence boundaries when combining Aleph Alpha’s disclosure with AA26-251A.

The advisory is therefore relevant because it demonstrates why upstream model lineage matters; it is not evidence of a Kolibri security incident.

Second-order capability provenance

Aleph Alpha’s direct training chain is documented: external teacher and judge models generate synthetic examples, regenerated answers, labels and judgments; Aleph Alpha then curates and validates those artifacts before using them in Kolibri’s training pipeline. AA26-251A describes a different, upstream process, alleging that some model developers acquired capabilities through unauthorized large-scale distillation from U.S. frontier models. When the two records are considered together, they imply a possible longer provenance chain in which capabilities originating in one frontier model could pass through an external teacher model, then into synthetic training material, and finally into Kolibri after Aleph Alpha’s own filtering and training stages.

That longer lineage is a possible provenance structure, not a demonstrated sample-by-sample chain. Neither Aleph Alpha’s documentation nor AA26-251A establishes that a particular Kolibri training example ultimately originated from a specific U.S. frontier model. The significance is conceptual: training Kolibri from random initialization establishes that its weights were independently optimized, but it does not establish that every reasoning pattern, response structure, factual representation or behavioral convention represented in its post-training data originated within Aleph Alpha.

Synthetic data therefore create a second-order capability channel. A model can be trained entirely from its own randomly initialized parameters while still learning from examples whose structure or content was produced by another model. In sovereignty terms, “trained from scratch in Europe” and “all capability provenance is European” are different claims. Kolibri strongly supports the first; the available evidence does not establish the second.

Values and biases are also supply-chain properties

Aleph Alpha explicitly recognizes that upstream models differ not only in capability but also in values and biases, and states that generator models are selected or avoided partly on that basis.165 The Kolibri model card similarly acknowledges that training material generated by Chinese models can transmit political bias and says Aleph Alpha applies filtering and dedicated alignment work to mitigate that risk.166

Synthetic-data provenance is therefore not merely an intellectual-property or legal-provenance issue. Teacher models can transmit characteristic factual errors, preferred solution patterns, answer formatting, refusal behavior, political or cultural priors, assumptions about institutions, security blind spots and recurring reasoning strategies. Subsequent SFT and reinforcement learning may modify, suppress or counteract these behaviors, but their upstream origin remains relevant when assessing how the final model was shaped.

This matters particularly in public-sector or regulated deployments, where sovereignty may be invoked not only to guarantee local execution and legal control but also to support control over institutional, cultural and political alignment. In that context, identifying which external models contributed to the training-information supply chain becomes part of the assurance problem itself.

Extraction is not poisoning

AA26-251A primarily concerns alleged unauthorized extraction and distillation of model capabilities. It should not be interpreted as evidence that the same companies deliberately inserted malicious examples into downstream training pipelines. Model extraction, data poisoning, backdoors and other adversarial-machine-learning mechanisms are distinct threat classes, a distinction also reflected in NIST’s adversarial-machine-learning taxonomy.167

That distinction matters technically for Kolibri. Aleph Alpha uses outputs generated by external models; it does not import the external models’ parameter tensors into Kolibri. A malicious or undesirable behavior embedded in an upstream checkpoint therefore would not automatically propagate as identical weights or an identical backdoor. Any influence would have to pass through generated examples, labels, judgments or other artifacts that survive Aleph Alpha’s curation and are subsequently consumed during training.

The relevant security concern is consequently output-mediated propagation. If an upstream teacher systematically generated biased, manipulated or trigger-conditioned examples, and if those examples survived filtering and were incorporated into the training mix, they could influence the downstream model. There is no evidence in the reviewed sources that this occurred in Kolibri. The architectural possibility is nevertheless sufficient to make upstream teacher provenance, validation and lineage legitimate supply-chain-security concerns.

From existing controls to supply-chain attestation

Kolibri’s Model Factory already provides several compensating controls. Aleph Alpha tracks intermediate and final checkpoints against application-relevant benchmark suites, performs grounding and abstention evaluations, maintains model, dataset and checkpoint lineage inside Savanna, and can regenerate problematic datasets before training successor checkpoints.

These controls improve the ability to detect broad regressions and remediate known problems, but they are not equivalent to complete provenance attestation. The public report does not establish a fully externally auditable sample-level record linking every generated example to an immutable teacher checkpoint or service endpoint.

For high-assurance supply-chain analysis, one would ideally know the exact teacher revision; whether generation occurred locally or through an API; provider and endpoint; inference runtime and quantization; prompt template; sampling configuration; generation date; source dataset; filtering stages; judge model; acceptance status; resulting dataset shard; and which downstream training runs consumed that shard.

NIST’s Generative AI Profile recommends provenance tracking for training and synthetic content and treats external generative-AI suppliers as third-party risk-management subjects.168 NIST’s adversarial-ML guidance recommends artifact-integrity mechanisms and cryptographic verification where appropriate, while ENISA treats data poisoning, adversarial attacks and ML-lifecycle security as explicit engineering concerns.169170

For sovereign AI, a model bill of materials should therefore extend beyond software libraries and accelerator drivers to include upstream teacher models. A useful release dossier would identify each material teacher, its exact revision where available, its role, approximate contribution, generation period, local-versus-hosted execution, license or service basis, validation mechanisms and relevant security advisories. The purpose is reverse lineage: if an upstream model is later found to contain a serious bias, provenance problem or compromise, the Model Factory should be able to determine which datasets, runs and released checkpoints were influenced by it.

The public documentation supports the following assessment.

Control objective Publicly documented Kolibri practice Residual assurance gap
Identification of principal SFT generators GLM-5.2, GLM-5.3 and Qwen3.8-27B are named contribution percentages are not disclosed
Recognition of teacher bias explicit discussion of model-specific values and biases residual source-specific bias after filtering is not independently quantified
Data-quality auditing poor completions and harmful datasets are audited and regenerated validation strength differs across datasets
Independent verification references, execution tests and judges are used in several pipelines not every generated example has independent semantic verification
Generator diversity multiple model families participate major SFT generation remains concentrated in GLM/Qwen families
Model lineage strong checkpoint and experiment lineage in Model Factory public sample-to-teacher lineage not demonstrated
Immutable teacher identification model names are published exact checkpoint hashes or API revisions are not publicly disclosed per sample
Generation access route not fully disclosed local weights versus direct API versus intermediary path often unresolved publicly
Artifact integrity Model Factory provides controlled datasets and runs cryptographically verifiable public synthetic-data manifests are not described
Political-bias mitigation filtering and dedicated alignment are documented source-specific residual measurements are not public
Teacher-specific backdoor testing broad security and behavioral evaluation exists source-specific trigger testing is not demonstrated
Supplier-advisory response no detailed public mechanism described unclear whether advisories automatically trigger lineage review and source ablation
Table 17: Kolibri’s published controls are substantial but do not amount to complete externally auditable synthetic-data supply-chain attestation.

The gap is therefore not between governance and no governance; it is between sophisticated internal production controls and externally auditable training-supply-chain lineage.

Source-ablation testing would strengthen sovereignty claims

Kolibri’s Model Factory is well suited to a control that is rarely discussed in sovereignty debates: source ablation. Smaller proxy models could be trained with and without specific teacher-derived subsets and then compared across capability metrics, political behavior, refusal patterns, factual-error clusters, solution similarity and security tests. A simple experimental design could compare a baseline dataset, the baseline plus GLM-derived material, the baseline plus Qwen-derived material, and the combined production mixture.

The purpose would not be to stigmatize a jurisdiction, but to measure how much observable model behavior is attributable to each external teacher family. This would turn provenance from a documentation problem into an empirical model-development question: rather than merely recording that a particular teacher contributed data, the Model Factory could estimate the marginal behavioral effect of that contribution. Kolibri’s existing proxy-training and mixture-search infrastructure makes this kind of experiment technically plausible.

Generator and judge independence

A second control follows from ordinary assurance engineering: the mechanism that generates an artifact should not automatically be treated as an independent authority for validating that same artifact. When one model generates synthetic data and another model from the same or a closely related family judges the output, correlated blind spots can survive the validation stage.

For high-assurance subsets, generator and verifier independence should therefore be increased wherever practical. Code should be executed rather than merely judged by another LLM when deterministic execution is available; mathematical outputs should be checked symbolically or numerically where feasible; public-administration facts should be compared with authoritative sources; and security-sensitive examples should use deterministic validators whenever such validators exist. A second LLM can still be useful as one layer of evaluation, but it should not automatically be treated as an independent source of truth.

Kolibri already applies variants of this principle selectively through execution-based checks, reference answers and separate judges. The sovereignty question is whether such independence can be made systematic for the datasets and domains most relevant to high-assurance or sovereign deployments.

Service provenance also matters

AA26-251A gives defensive recommendations to providers whose models are targeted for distillation.171 That has a separate implication for any synthetic-data factory that relies on hosted external models: a model name is not necessarily a stable or reproducible artifact. A provider can update the underlying model, alter system prompts, change request routing, modify safety policies or move traffic to a different serving configuration without changing the externally visible product name.

There is no evidence that such a mechanism affected Aleph Alpha. The architectural implication is narrower: reliable provenance requires more than recording GLM-5.2 or Qwen3.8-27B. The access path, concrete model revision and serving conditions also matter. Local execution of retained open weights can reduce this ambiguity because the exact artifact can be hashed, archived and rerun under a known inference configuration. Hosted APIs generally require additional provider-side attestation if the downstream organization wants equivalent provenance guarantees.

Reproductive sovereignty is stricter than pipeline ownership

This exposes a distinction in the phrase we own the entire pipeline.

Aleph Alpha clearly controls the orchestration of Kolibri’s production pipeline: what is generated, which data are accepted, how datasets are mixed, which experiments are run and how final weights are produced.

But consider a stronger test:

Could Aleph Alpha reproduce a functionally equivalent post-training corpus if it permanently lost access to GLM, Qwen and every other non-European generator?

The public documentation does not establish that it could do so without degradation. The relevant dimensions therefore differ.

Dimension Kolibri
control of the released model weights strong
independent inference after download strong
control over architecture and tokenizer strong
control over training and evaluation strong
control over synthetic-data selection and filtering strong
independent origin of all synthetic knowledge inputs no
ability to regenerate equivalent synthetic data without external teachers not established
independent European accelerator supply no
public reproducibility of the complete Model Factory incomplete
Table 18: Control of a pipeline is not identical to independence from every upstream input.

This is not unusual. Industrial systems routinely rely on external inputs while retaining strong control over the process. The relevant sovereignty criterion is therefore not autarky but substitutability.

An external component becomes strategically dangerous when its loss creates an unacceptable interruption and no timely substitute exists.

Mistral Large 4 changes the substitutability question

This is where the Mistral Large 4 release changes the Kolibri appendix most interestingly.

Aleph Alpha already used a Mistral model, Mistral-Nemo, in part of its German pre-training-data pipeline. Before ML4, the most capable disclosed synthetic-data teachers in Kolibri’s post-training pipeline were Chinese models such as GLM and Qwen.

Mistral Large 4 now gives Europe a much more capable indigenous candidate teacher family. Its Artificial Analysis score of 38 is below GLM-5.3 Max at 45 and Kimi K3 Max at 44, but it is much closer to that frontier class than previous European models.172

That changes the option set for future European training pipelines. It does not establish that ML4 can replace GLM-5.3 or Qwen3.8-27B in every task for which Aleph Alpha used them. Teacher suitability is task-specific, and Mistral has not published a controlled substitution experiment against Kolibri’s SFT generators.

The defensible conclusion is narrower:

Mistral Large 4 materially improves the plausibility of European substitution for some external teacher-model roles, but substitutability has not yet been demonstrated for Kolibri’s specific data-generation workloads.

There is also a timing issue. As of 7 October, the ML4 public weights are not yet available. API access exists, and verified partners may receive less-moderated access, but a reproducible local teacher artifact for ordinary third-party use remains a future state until the announced weight release occurs.173

If those weights are released as announced, the sovereignty effect will be larger: a European model factory could potentially retain an exact local ML4 artifact rather than rely on a remote non-European API for some synthetic-data tasks.

Public-sector implications

For a ministry, regulator, municipality or critical-infrastructure operator, Kolibri’s use of Chinese models does not make a self-hosted Kolibri deployment non-sovereign in the operational sense.

Those teachers are not in the runtime path once Kolibri has been trained and downloaded. A self-hosted Kolibri deployment can keep prompts, retrieved documents and outputs inside infrastructure chosen by the customer.

Losing future access to GLM, Qwen or Kimi would not by itself erase information already incorporated into the released Kolibri checkpoint. Runtime sovereignty and training-information provenance are therefore different concerns. A public-sector buyer evaluating Kolibri should distinguish at least the following dimensions.

Sovereignty dimension Assessment
Runtime sovereignty strong when self-hosted
Final-artifact control strong
Model-production control strong
Training-information provenance mixed
Upstream-teacher substitutability incomplete / not publicly demonstrated
Compute-infrastructure control substantial operationally, incomplete physically
Accelerator sovereignty weak
Supply-chain auditability unusually detailed for a commercial model, but incomplete at immutable teacher/sample level
Evidence of compromise from AA26-251A none established
Corporate continuity future state depends partly on the announced Cohere combination
Table 19: Sovereignty dimensions relevant to a public-sector Kolibri procurement.

For higher-assurance procurement, the appropriate response is therefore a provenance dossier rather than a binary declaration that a model is either European or not European.

Such a dossier should ideally explain which material external teachers were used, their concrete revisions, approximate contributions, generation periods and access mechanisms; how generated data were filtered and validated; whether security advisories affected any supplier; whether reverse lineage is available; and how affected material could be removed and regenerated if an upstream problem were later discovered.

The same logic should apply to Mistral Large 4 after the public weight release. European model should not substitute for a technical dossier on training data, dependencies, artifact integrity and the infrastructure required to reproduce the next generation.

Toward a stricter definition of sovereign foundation models

The comparison exposes three recurring external dependencies: physical dependence on accelerators and semiconductor supply; informational dependence on external training-data sources or teacher models; and organizational dependence on the company or institution that retains the engineers, code, datasets, compute allocations and decision rights. Their failure modes differ: loss of accelerators can block a new training run, loss of a teacher can prevent regeneration of an equivalent synthetic corpus, and corporate change can move IP or production decisions. A retained model checkpoint may survive all three.

This is why a sovereign model artifact and a sovereign model-production system are different objects. The European comparison suggests a hierarchy rather than a binary label:

  • At the weakest level is jurisdictional consumption: a model is accessed under a European contract or through an EU-located service.
  • Above that is deployment sovereignty: the customer can retain and execute the model independently.
  • Above that is model sovereignty: the model artifact can be inspected, adapted and preserved.
  • Above that is producer sovereignty: a European organization controls the architecture, tokenizer, data pipeline, training, post-training and evaluation necessary to produce the model.
  • Above that is reproductive sovereignty: the organization can regenerate future models and material training inputs without depending on irreplaceable external actors.
  • Finally, full-stack industrial sovereignty would require acceptable control or timely substitutability across compute systems, accelerators, high-bandwidth memory, semiconductor manufacturing, packaging, networking and the software stack needed to operate them.

Mistral Large 4 pushes Europe materially upward on producer sovereignty and infrastructure control. Kolibri provides unusually strong evidence for model-production traceability and currently exercisable deployment sovereignty. Apertus provides the strongest evidence for public reproductive transparency. None of them demonstrates full-stack semiconductor sovereignty.

Conclusion

Mistral Large 4 changes the European AI baseline in a material way. Its Artificial Analysis score of 38, trillion-parameter sparse architecture and training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres demonstrate that Europe now possesses a private model developer operating much closer to the global frontier while retaining direct control of the datacentre layer.174175

The release does not settle the sovereignty question. As of 7 October 2026, ML4 remains a preview: its public weights and final weight license are still pending, and its accelerator substrate remains non-European. Kolibri exposes a different limitation: a European-controlled model-production pipeline can still depend on non-European teacher models for synthetic training information.

Taken together, the two projects sharpen rather than resolve the concept of sovereign AI. Mistral demonstrates that European producer capability and infrastructure control can approach frontier scale; Kolibri demonstrates why training-information provenance and reproducibility must also be included in the assessment. Apertus, EuroLLM, Salamandra, Teuken and OpenEuroLLM add further strengths in public reproducibility, multilingual coverage and institutional depth.

The resulting European position is therefore best described as an emerging portfolio of sovereignty capabilities, not a completed sovereign stack. Europe can build serious foundation models, operate large training infrastructure and release independently deployable artifacts. The unresolved test is whether the critical external layers—accelerators, upstream teacher models and organizational dependencies—are sufficiently substitutable that losing one of them would not prevent the next model generation.

See also machine learning longforms

See also posts

Back to top

Footnotes

  1. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  2. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  3. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  4. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  5. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  6. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  7. Barlow, M. (2026). Model Training as Code. Aleph Alpha Research, 22 May 2026. Describes Savanna, its CI model, immutable artefacts, lineage, workflow engine and collaboration model. Technical article↩︎

  8. Barlow, M. (2026). Model Training as Code. Aleph Alpha Research, 22 May 2026. Describes Savanna, its CI model, immutable artefacts, lineage, workflow engine and collaboration model. Technical article↩︎

  9. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  10. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  11. Barlow, M. (2026). Model Training as Code. Aleph Alpha Research, 22 May 2026. Describes Savanna, its CI model, immutable artefacts, lineage, workflow engine and collaboration model. Technical article↩︎

  12. Barlow, M. (2026). Model Training as Code. Aleph Alpha Research, 22 May 2026. Describes Savanna, its CI model, immutable artefacts, lineage, workflow engine and collaboration model. Technical article↩︎

  13. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  14. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  15. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  16. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  17. Sassoon, J. (2026). Scaling Pre-Training in Practice: A Hierarchical Approach. Aleph Alpha Research, 30 September 2026. Documents a 30B-A3B scaling study from 16 to 512 B200 GPUs, culminating in 35.3% MFU with FSDP128/DP4 and a 6% per-GPU efficiency loss from the 16-GPU optimum. Technical article↩︎

  18. Sassoon, J. (2026). Scaling Pre-Training in Practice: A Hierarchical Approach. Aleph Alpha Research, 30 September 2026. Documents a 30B-A3B scaling study from 16 to 512 B200 GPUs, culminating in 35.3% MFU with FSDP128/DP4 and a 6% per-GPU efficiency loss from the 16-GPU optimum. Technical article↩︎

  19. Sassoon, J. (2026). Scaling Pre-Training in Practice: A Hierarchical Approach. Aleph Alpha Research, 30 September 2026. Documents a 30B-A3B scaling study from 16 to 512 B200 GPUs, culminating in 35.3% MFU with FSDP128/DP4 and a 6% per-GPU efficiency loss from the 16-GPU optimum. Technical article↩︎

  20. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  21. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  22. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  23. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  24. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  25. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  26. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  27. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  28. Hartel, A., & Dürr, N. (2026). Sauerkraut, Not Burgers: Why German LLMs Need German Data. Aleph Alpha Research, 8 September 2026. Describes the German data pipeline, insufficiency of open German datasets, language-specific filtering and the strategy of organic data plus synthetic rephrasing. Technical article↩︎

  29. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  30. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  31. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  32. Finken, N., & Thel, S. (2026). Through the Valley of Tears: Cold Starting German Reasoning in LLMs. Aleph Alpha Research, 24 September 2026. Documents the approximately 796k-sample German SFT pipeline, reasoning-prefill distillation, language-consistency measurements and the observed repetition-loop failure mode. Technical article↩︎

  33. Finken, N., & Thel, S. (2026). Through the Valley of Tears: Cold Starting German Reasoning in LLMs. Aleph Alpha Research, 24 September 2026. Documents the approximately 796k-sample German SFT pipeline, reasoning-prefill distillation, language-consistency measurements and the observed repetition-loop failure mode. Technical article↩︎

  34. Finken, N., & Thel, S. (2026). Through the Valley of Tears: Cold Starting German Reasoning in LLMs. Aleph Alpha Research, 24 September 2026. Documents the approximately 796k-sample German SFT pipeline, reasoning-prefill distillation, language-consistency measurements and the observed repetition-loop failure mode. Technical article↩︎

  35. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  36. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  37. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  38. Maskey, S., & Wirges, S. (2026). Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack. Aleph Alpha Research, 10 September 2026. Shows a controlled ranking reversal among pre-training checkpoints after identical downstream training and argues that checkpoint selection must consider later stages. Technical article↩︎

  39. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  40. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  41. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  42. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  43. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  44. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  45. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  46. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  47. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  48. Boll, B. (2026). Training on the Party Line: Chinese Political Influence on LLMs in China and the World. Aleph Alpha Research, 28 September 2026. Describes Aleph Alpha’s evaluation of political bias in Chinese models and synthetic datasets and states that the company screens training data, adds targeted alignment data and evaluates political alignment. Technical article↩︎

  49. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  50. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  51. Parcalabescu, L. (2026). Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models. Aleph Alpha Research, 8 August 2026. Describes the grounding/abstention training mechanism used in Kolibri’s post-training environments. Technical article↩︎

  52. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  53. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  54. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  55. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  56. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  57. Körner, S. (2026). Transparency as One Pillar of Sovereign AI: We Signed the Code of Practice on Transparency of AI-generated Content. Aleph Alpha, 4 August 2026. Describes Aleph Alpha’s signature of the EU transparency code and its stated position on provenance and machine-readable transparency. Official statement↩︎

  58. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  59. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  60. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for exact parameter counts, context guidance, FP8 format, hardware requirements, training resources, data-curation summary, post-training configuration, risks, public training-content summary and license boundary. Official model card↩︎

  61. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  62. Barlow, M. (2026). Model Training as Code. Aleph Alpha Research, 22 May 2026. Describes Savanna, its CI model, immutable artefacts, lineage, workflow engine and collaboration model. Technical article↩︎

  63. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  64. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  65. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Primary source for the release, sovereignty framing, Kolibri-Origin comparison, production chronology, architecture trade-offs, >200T raw-token processing claim, automatic run recovery, German-data summary, post-training totals and public deployment position. Official announcement↩︎

  66. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  67. Aleph Alpha and Cohere. (2026). Cohere and Aleph Alpha Sign Agreement to Become the First Transatlantic Sovereign AI Solution. 16 September 2026. The announced transaction remains subject to final regulatory approvals and provides for a combined company operating globally as Cohere with Germany- and Canada-based headquarters, R&D and leadership. Official announcement↩︎

  68. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  69. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for architecture, optimization, data construction, distributed training, SFT/RL details, evaluation and Model Factory internals. Official technical report↩︎

  70. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  71. Artificial Analysis. (2026). Artificial Analysis Intelligence Index v4.3.2. Current index comprising ten professional, agentic, scientific, knowledge and long-context evaluations. Accessed 7 October 2026. Index methodology↩︎

  72. Artificial Analysis. (2026). Model evaluation pages used for the bridge calibration. Current values as accessed 7 October 2026 for GLM-4.7 Flash, Nemotron 3 Nano, Qwen3.5 35B-A3B, Qwen3.6 35B-A3B, Qwen3-Next 80B-A3B Thinking, Gemma 4 26B-A4B, GPT-OSS 120B, Mistral Small 4, GLM-4.5 Air, Nemotron 3 Super and Qwen3.8 27B. Artificial Analysis marks some of these values as estimated; Table Table 6 preserves that distinction. Artificial Analysis models↩︎

  73. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  74. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  75. Artificial Analysis. (2026). GPT-6 Astra (Max). Intelligence Index 53; Humanity’s Last Exam 55%; AA-Omniscience 43; AA-LCR 81%. Accessed 7 October 2026. Model evaluation↩︎

  76. Artificial Analysis. (2026). Claude Opus 5.5 (Max, Default Fallback). Intelligence Index 58; Humanity’s Last Exam 61%; AA-Omniscience 46; AA-LCR 85%. Accessed 7 October 2026. Model evaluation↩︎

  77. Artificial Analysis. (2026). Artificial Analysis Intelligence Index v4.3.2. Current index comprising ten professional, agentic, scientific, knowledge and long-context evaluations. Accessed 7 October 2026. Index methodology↩︎

  78. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  79. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  80. Artificial Analysis. (2026). Qwen3.8 27B. The xhigh reasoning configuration scores 34 on the Intelligence Index, including 34% on Humanity’s Last Exam, approximately −10 on AA-Omniscience and 82% on AA-LCR. Accessed 7 October 2026. Model evaluation↩︎

  81. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  82. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  83. Artificial Analysis. (2026). GPT-6 Astra (Max). Intelligence Index 53; Humanity’s Last Exam 55%; AA-Omniscience 43; AA-LCR 81%. Accessed 7 October 2026. Model evaluation↩︎

  84. Artificial Analysis. (2026). Claude Opus 5.5 (Max, Default Fallback). Intelligence Index 58; Humanity’s Last Exam 61%; AA-Omniscience 46; AA-LCR 85%. Accessed 7 October 2026. Model evaluation↩︎

  85. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  86. Artificial Analysis. (2026). Artificial Analysis Intelligence Index v4.3.2. Current index comprising ten professional, agentic, scientific, knowledge and long-context evaluations. Accessed 7 October 2026. Index methodology↩︎

  87. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  88. Artificial Analysis. (2026). Model evaluation pages used for the bridge calibration. Current values as accessed 7 October 2026 for GLM-4.7 Flash, Nemotron 3 Nano, Qwen3.5 35B-A3B, Qwen3.6 35B-A3B, Qwen3-Next 80B-A3B Thinking, Gemma 4 26B-A4B, GPT-OSS 120B, Mistral Small 4, GLM-4.5 Air, Nemotron 3 Super and Qwen3.8 27B. Artificial Analysis marks some of these values as estimated; Table Table 6 preserves that distinction. Artificial Analysis models↩︎

  89. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  90. Artificial Analysis. (2026). Qwen3.8 27B. The xhigh reasoning configuration scores 34 on the Intelligence Index, including 34% on Humanity’s Last Exam, approximately −10 on AA-Omniscience and 82% on AA-LCR. Accessed 7 October 2026. Model evaluation↩︎

  91. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  92. Artificial Analysis. (2026). GPT-6 Astra (Max). Intelligence Index 53; Humanity’s Last Exam 55%; AA-Omniscience 43; AA-LCR 81%. Accessed 7 October 2026. Model evaluation↩︎

  93. Artificial Analysis. (2026). Claude Opus 5.5 (Max, Default Fallback). Intelligence Index 58; Humanity’s Last Exam 61%; AA-Omniscience 46; AA-LCR 85%. Accessed 7 October 2026. Model evaluation↩︎

  94. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  95. Artificial Analysis. (2026). Qwen3.8 27B. The xhigh reasoning configuration scores 34 on the Intelligence Index, including 34% on Humanity’s Last Exam, approximately −10 on AA-Omniscience and 82% on AA-LCR. Accessed 7 October 2026. Model evaluation↩︎

  96. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  97. Artificial Analysis. (2026). GPT-6 Astra (Max). Intelligence Index 53; Humanity’s Last Exam 55%; AA-Omniscience 43; AA-LCR 81%. Accessed 7 October 2026. Model evaluation↩︎

  98. Artificial Analysis. (2026). Claude Opus 5.5 (Max, Default Fallback). Intelligence Index 58; Humanity’s Last Exam 61%; AA-Omniscience 46; AA-LCR 85%. Accessed 7 October 2026. Model evaluation↩︎

  99. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  100. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  101. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  102. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  103. Artificial Analysis. (2026). AA-Omniscience. Evaluation rewarding correct answers, penalizing incorrect answers and assigning no penalty to abstention; the index ranges from −100 to +100. Accessed 7 October 2026. Artificial Analysis↩︎

  104. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  105. Artificial Analysis. (2026). Qwen3.8 27B. The xhigh reasoning configuration scores 34 on the Intelligence Index, including 34% on Humanity’s Last Exam, approximately −10 on AA-Omniscience and 82% on AA-LCR. Accessed 7 October 2026. Model evaluation↩︎

  106. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  107. Artificial Analysis. (2026). GPT-6 Astra (Max). Intelligence Index 53; Humanity’s Last Exam 55%; AA-Omniscience 43; AA-LCR 81%. Accessed 7 October 2026. Model evaluation↩︎

  108. Artificial Analysis. (2026). Claude Opus 5.5 (Max, Default Fallback). Intelligence Index 58; Humanity’s Last Exam 61%; AA-Omniscience 46; AA-LCR 85%. Accessed 7 October 2026. Model evaluation↩︎

  109. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  110. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  111. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  112. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  113. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  114. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  115. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  116. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  117. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  118. Artificial Analysis. (2026). Mistral model evaluations. As accessed 7 October 2026, Mistral Large 4 Preview scores 38, Mistral Medium 3.5 scores 14, Mistral Small 4 scores 11 and Mistral Large 3 scores 9 on the current Intelligence Index. Mistral Large 4 was still an API preview at that date; its public weight release was announced for later in October. These values are time-sensitive because Artificial Analysis updates model evaluations and index versions. Mistral models↩︎

  119. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  120. Arena. (2026). Text Arena overall leaderboard. 2 October 2026 snapshot. Gemini 4 Argon High ranked first, Claude Opus 5.5 High fourth and Kimi K3 Max sixteenth. Leaderboard↩︎

  121. Arena. (2026). Open-weight text leaderboard. 11 September 2026 snapshot, with GLM-5.3 Max and Kimi K3 Max in the leading two open-weight positions. Open-weight leaderboard↩︎

  122. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Official launch announcement and selected benchmark/serving comparison. Official announcement↩︎

  123. Artificial Analysis. (2026). Qwen3.8 27B. The xhigh reasoning configuration scores 34 on the Intelligence Index, including 34% on Humanity’s Last Exam, approximately −10 on AA-Omniscience and 82% on AA-LCR. Accessed 7 October 2026. Model evaluation↩︎

  124. Artificial Analysis. (2026). Qwen3.8 27B versus Qwen3.8 2.4T-A95B. The larger Qwen release scores approximately 40 on the Intelligence Index and has 2.4T total / 95B active parameters. Accessed 7 October 2026. Comparison↩︎

  125. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  126. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  127. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  128. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  129. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for the complete post-training benchmark tables, baseline deployment settings, model architecture and training details. Technical report↩︎

  130. Artificial Analysis. (2026). GLM-5.3 Max. Intelligence Index 45; 753B total and 40B active parameters; 1M-token context; downloadable weights under the GLM-5.3 License. Accessed 7 October 2026. Model evaluation↩︎

  131. Artificial Analysis. (2026). Kimi K3 Max. Intelligence Index 44; 2.8T total and 104B active parameters; 1M-token context; downloadable weights under the Kimi K3 License. Accessed 7 October 2026. Model evaluation↩︎

  132. Artificial Analysis. (2026). GPT-6 Astra (Max). Intelligence Index 53; Humanity’s Last Exam 55%; AA-Omniscience 43; AA-LCR 81%. Accessed 7 October 2026. Model evaluation↩︎

  133. Artificial Analysis. (2026). Claude Opus 5.5 (Max, Default Fallback). Intelligence Index 58; Humanity’s Last Exam 61%; AA-Omniscience 46; AA-LCR 85%. Accessed 7 October 2026. Model evaluation↩︎

  134. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  135. Mistral AI. (2026). Mistral Large 4 — model documentation. Reports 1.05T total parameters, 49B active parameters, a 1.6B vision encoder, multimodal input and a documented 1M-token context. Official documentation↩︎

  136. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  137. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for model architecture, training pipeline, SFT dataset generation, external teacher models, evaluation and Model Factory. Technical report↩︎

  138. Cybersecurity and Infrastructure Security Agency, National Security Agency, and Federal Bureau of Investigation. (2026). AA26-251A: China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies. 8 September 2026. The advisory contains U.S. government allegations concerning Z.AI, Alibaba, Moonshot AI and other companies; those allegations should not be extended to exact downstream checkpoints without separate evidence. CISA advisory↩︎

  139. Swiss AI Initiative. (2025). Apertus-70B model card and training resources. The model card reports 15T pre-training tokens, 4,096 NVIDIA GH200 GPUs and public reconstruction scripts, training framework and intermediate checkpoints. Model card↩︎

  140. Mistral AI. (2026). Mistral Large 4 — model documentation. Reports 1.05T total parameters, 49B active parameters, a 1.6B vision encoder, multimodal input and a documented 1M-token context. Official documentation↩︎

  141. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  142. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  143. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  144. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  145. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  146. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  147. Mistral AI. (2026). Mistral-Large-4.0-1T05-A52B — Upcoming release. Mistral’s Hugging Face release page identifies the planned public checkpoint and states that 52B parameters are active when embeddings and output layers are included. Upcoming release↩︎

  148. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  149. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  150. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  151. EuroLLM Team. (2025). EuroLLM-22B. Reports approximately 4T training tokens and training on 400 NVIDIA H100 GPUs on MareNostrum 5 under EuroHPC access. Project article↩︎

  152. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  153. Barcelona Supercomputing Center Language Technologies Unit. Salamandra model card. Documents the 2B/7B/40B family, 12.875T-token multilingual pre-training corpus, 35 European languages, Apache 2.0 licensing, and public training scripts/configurations. Model card↩︎

  154. OpenGPT-X / Fraunhofer. Teuken model cards and project documentation. Current v0.6 model cards describe a 7B multilingual model trained on 6T tokens in all 24 official EU languages; earlier v0.4 commercial variants were released under Apache 2.0. Current model card↩︎

  155. OpenEuroLLM. OpenEuroLLM. EU-funded programme for openly developed European foundation models, datasets, tools and training infrastructure. Official project↩︎

  156. OpenEuroLLM. (2025). Strategic access to EuroHPC resources granted to OpenEuroLLM. Reports more than ten million GPU-hours across LUMI, Leonardo, JUPITER and MareNostrum 5. Project announcement↩︎

  157. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  158. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  159. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for model architecture, training pipeline, SFT dataset generation, external teacher models, evaluation and Model Factory. Technical report↩︎

  160. Aleph Alpha. (2026). Kolibri Has Landed: A Sovereign Open-Weight Model. 3 October 2026. Reports 174B generated synthetic SFT tokens and an approximately 268B-token filtered SFT mixture. Official announcement↩︎

  161. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for model architecture, training pipeline, SFT dataset generation, external teacher models, evaluation and Model Factory. Technical report↩︎

  162. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for model architecture, training pipeline, SFT dataset generation, external teacher models, evaluation and Model Factory. Technical report↩︎

  163. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for model architecture, training pipeline, SFT dataset generation, external teacher models, evaluation and Model Factory. Technical report↩︎

  164. Cybersecurity and Infrastructure Security Agency, National Security Agency, and Federal Bureau of Investigation. (2026). AA26-251A: China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies. 8 September 2026. The advisory contains U.S. government allegations concerning Z.AI, Alibaba, Moonshot AI and other companies; those allegations should not be extended to exact downstream checkpoints without separate evidence. CISA advisory↩︎

  165. Aleph Alpha. (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Primary source for model architecture, training pipeline, SFT dataset generation, external teacher models, evaluation and Model Factory. Technical report↩︎

  166. Aleph Alpha. (2026). Kolibri-1 model card. Primary source for release configuration, licensing, intended use and discussion of political-bias risks in model-generated training material. Model card↩︎

  167. National Institute of Standards and Technology. (2025). Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. NIST AI 100-2e2025. NIST publication↩︎

  168. National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. NIST publication↩︎

  169. National Institute of Standards and Technology. (2025). Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. NIST AI 100-2e2025. NIST publication↩︎

  170. European Union Agency for Cybersecurity. Securing Machine Learning Algorithms. ENISA publication↩︎

  171. Cybersecurity and Infrastructure Security Agency, National Security Agency, and Federal Bureau of Investigation. (2026). AA26-251A: China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies. 8 September 2026. The advisory contains U.S. government allegations concerning Z.AI, Alibaba, Moonshot AI and other companies; those allegations should not be extended to exact downstream checkpoints without separate evidence. CISA advisory↩︎

  172. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎

  173. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  174. Mistral AI. (2026). Introducing Mistral Large 4. 6 October 2026. Primary source for the public-preview status, 1T/49B-active headline architecture, training on 3,800 NVIDIA Grace Blackwell GPUs in Mistral-owned European datacentres, 160+ languages, European service operation, benchmark claims and large-scale RL infrastructure. Official announcement↩︎

  175. Artificial Analysis. (2026). Mistral Large 4 Preview — Intelligence, Performance & Price Analysis. Accessed 7 October 2026. Reports an Intelligence Index of 38, 116.1 output tokens/s on Mistral’s API, current 524k served context in the tested endpoint and the preview’s present non-public-weight status. Artificial Analysis separately characterizes ML4 as the most capable model from outside the U.S. and China in its current index. Model evaluation↩︎