Explainer

The open-weight model landscape, September 2026: what changed and what it means for sovereign deployment

DeepSeek moved to V4, Llama moved to 4.5, and reasoning stopped being a separate model family. Here is what actually shifted, and why it matters if you are choosing what to run on your own hardware.

The short answer

The open-weight field moved substantially in 2026. DeepSeek's V3 and R1 generation, a separate base model and a separate reasoning model, has been superseded by V4: a single mixture-of-experts family with thinking and non-thinking modes built in, and a planned separate R2 reasoning model was folded into that release instead of shipping on its own. Meta's Llama 4 line was extended with Llama 4.5 in June 2026, the same architecture and hardware class as Llama 4. For anyone evaluating what to self-host, the practical shift is that a single current-generation model now does what used to need two, which changes both the hardware calculus and the vendor evaluation questions worth asking.

Model releases move fast enough that an article written in the spring can misdescribe the field by autumn. Two changes since the summer are worth setting out plainly, because they affect what a genuine evaluation of open-weight models should now ask.

DeepSeek collapsed two models into one

DeepSeek's previous generation split cleanly in two: V3 as the general-purpose base model, R1 as a separate reasoning-focused model trained on top of it. Evaluating DeepSeek meant picking which of the two fit your task, and a dedicated successor to R1, referred to as R2, was widely expected to continue that split.

It did not ship that way. DeepSeek's current generation, V4, is a single mixture-of-experts model, with V4-Pro as the general-availability release and V4.1-Flash following in September 2026 as a faster variant tuned for coding-agent workloads. Reasoning is now a mode within the one model rather than a separate weights file to download and swap to. For anyone who built an evaluation process around "which of these two models do I need," that process needs revisiting: the current question is closer to "does this one model's thinking mode do what I need," which is a different, and generally simpler, deployment decision.

Llama's line extended, not replaced

Meta's Llama 4 family gained a Llama 4.5 release in June 2026. This was an extension within the same architecture and hardware class as Llama 4, not a redesign, so hardware sizing guidance written for Llama 4 largely still applies to 4.5. That continuity is itself worth noting: not every open-weight update reshuffles what you need to run it, and Llama's approach here has been more incremental than DeepSeek's architectural consolidation.

What this means for evaluation, not just for spec sheets

Neither of these changes is really about which model wins a benchmark chart. They are about what changes in a genuine sovereign or air-gapped evaluation:

  • Fewer separate weight files to manage. A consolidated model such as DeepSeek V4 reduces the operational overhead of maintaining, updating and verifying two separate downloaded artefacts where one used to do the job of two.
  • Total parameter count still is not the sizing number that matters. Both DeepSeek V4 and Llama 4's largest variants are mixture-of-experts designs, activating a fraction of their total parameters per request while still needing the full weight set resident in memory somewhere. A model's headline parameter figure tells you almost nothing about what hardware it needs without also knowing its active-parameter count and quantisation.
  • "Reasoning model" as a separate procurement category is fading. If your evaluation criteria still ask "do we need the base model or the reasoning model," that question is becoming the wrong shape for at least this family. The current, more useful question is whether a model's built-in thinking mode is worth the extra latency for your specific task.

None of this changes the fundamentals of evaluating a model for a genuinely sovereign, offline deployment: where the weights physically sit, whether inference can run with no network path out, and whether the licence permits your intended use are still the questions that separate sovereignty from private-cloud marketing with a sovereign label attached. What has changed is the shape of the model landscape those questions get applied to. Our buyer's checklist for sovereign AI vendors covers that evaluation in full; this piece is about keeping the underlying model landscape current, not about restating it.

Disclosure

Unified is published by Mickai LTD, which builds in this field. For Mickai's own approach, see mickai.co.uk.

Questions readers ask

Is DeepSeek R2 available?
No. A dedicated R2 reasoning model was reported in development but was not released as a separate model. DeepSeek's current generation, V4, includes thinking and non-thinking modes in one model instead, which is the practical successor to the old base-model-plus-reasoning-model split.
Is DeepSeek V3 still worth using?
It still runs and its weights are still published, but it is a superseded generation. V4-Pro is the current flagship, with V4.1-Flash following in September 2026 as a faster, coding-optimised variant. Anything written or benchmarked against V3 should be treated as describing last year's model, not the current one.
Does a bigger open-weight model automatically mean better for a sovereign deployment?
No. Total parameter count is a poor proxy for what you can actually run. Mixture-of-experts models such as DeepSeek's and Llama 4's activate only a fraction of their total parameters per request, but the full weight set still has to fit in memory somewhere. The practical question for a sovereign deployment is what fits your actual hardware at an acceptable speed, not which model has the largest headline parameter count.
MICKAI®

Published by Mickai LTD. Written by the Mickai team.

Unified covers the field broadly and treats Mickai as one example within it. About the journal and the team.

Mickarle Wagstaff-Irons - Micky Irons, full name Mickarle Sean Junior Wagstaff-Irons. Founder and Interim CEO of Mickai. Biography and related work.

Keep reading

  • Landscape

    Who builds sovereign AI in the UK?

    Sovereign AI is a real field with serious participants. Here is how to think about the landscape, and where a British company is pushing past the baseline.

    2 min read

  • Explainer

    What is sovereign AI? A plain-English definition

    Sovereign AI means running artificial intelligence entirely under your own control, on infrastructure you own, with nothing forced through an outside provider. Here is what that means in practice, and why regulated organisations are moving toward it.

    5 min read