← Back to blog
EN | NO
AI Governance 10 min read

Kimi K3 and the open-weight wave: is American AI actually safer than Chinese?

Uros Vujic 23. juli 2026

A model with 2.8 trillion parameters, and a question nobody likes to answer precisely

On July 16, 2026, Chinese company Moonshot AI launched Kimi K3 – 2.8 trillion parameters, the largest open-weight AI model ever released. The weights only become publicly available on July 27, under a modified MIT license, so what exists right now is API access and the company's own benchmark numbers. Those should be taken with a grain of salt until someone outside Moonshot has actually tested it. But the launch itself triggered something: Asian markets fell sharply the same day, in what several outlets called a new "DeepSeek moment" – a repeat of the shock from January 2025, when the market first understood that Chinese labs could match American frontier models at a fraction of the cost.

Prices are already moving. Kimi K3 costs roughly $3 per million input tokens and $15 per million output tokens – markedly cheaper than American frontier models. Anthropic cut Sonnet pricing on June 30. OpenAI's Sam Altman has signaled they could go down to a quarter of current pricing if needed. Google cut its Ultra subscription price by 20 percent in May. And on OpenRouter, one of the largest AI model marketplaces, the share of tokens routed to Chinese open models went from roughly 11 percent a year ago to as much as 46 percent this July.

That's one story: open models are pushing prices down across the entire industry, regardless of who you actually use.

The other story is harder, and it's the one this article is actually about: should a Norwegian business be more nervous about running a Chinese open-weight model than an American one? And is "American" actually a safe default – or do we trust it out of habit, not documented reason?


What's actually documented about Chinese models

Let's start with the uncomfortable part, because it's real and it helps no one to pretend otherwise.

Political censorship in Chinese models is well-documented and reproducible. Independent tests consistently find that DeepSeek and similar models refuse to answer, or answer with the regime's official framing, on questions about Tiananmen Square 1989, Taiwan's status, Xinjiang, and Hong Kong. This isn't an isolated finding – it's been reproduced in academic studies and in a US congressional committee report. It's a real, measurable property of the model, not rumor.

There's also a newer and more uncomfortable data point: the US security body CAISI, under NIST, tested DeepSeek and found the model roughly 12 times more likely than American frontier models to comply with malicious instructions, with pro-Beijing narrative bias built in. A separate red-teaming firm, Enkrypt AI, found DeepSeek-R1 eleven times more likely to generate harmful content than OpenAI's o1 model. These aren't numbers we should wave away just to land on a neat symmetrical conclusion. They're real, and they're larger than the equivalent weaknesses documented in American open models like Llama.

Norway's own National Security Authority (NSM) explicitly warned against DeepSeek, and the Data Protection Authority (Datatilsynet) has flagged GDPR concerns. Both warnings center on China's Intelligence Law – specifically Article 7, which requires Chinese companies to cooperate with state intelligence services on request. This isn't a theoretical concern; it's written into Chinese law.


What's less well known about American models

Here the picture gets more complicated than "American = safe."

The CLOUD Act gives US authorities the right to compel American companies to hand over data regardless of where in the world it's physically stored – including in the EU. The European Data Protection Board (EDPB) has explicitly flagged this as a conflict with GDPR following the Schrems II ruling. Structurally, this is the closest American parallel to China's Intelligence Law: both give the state a legal route into data held by a company under their jurisdiction, regardless of where the customer sits.

And then there's something that surprised us in the research: in February 2026, Anthropic was reportedly blacklisted by the Trump administration after refusing to grant the Pentagon unrestricted "lawful purpose" access for mass surveillance and autonomous weapons systems. In June the same year, foreign access to two of Anthropic's models was reportedly suspended by US authorities, and the company had to withdraw access entirely. The point here isn't that Anthropic did anything wrong – quite the opposite, it's a courageous decision. The point is that even an American company isn't exempt from its own state placing political conditions on non-American customers' access to the product. It isn't just China that has a state-access problem. The US has its own, through a different mechanism.


The distinction that actually determines the risk: hosted app or your own weights

Here's the point that gets lost in most of the debate: there's a big difference between using DeepSeek's app or API, and downloading DeepSeek's open weights and running them on your own or European infrastructure.

When you use the hosted service, your prompts go to servers the Chinese company controls – and the intelligence law applies directly. That's exactly what NSM and Datatilsynet warned against. When you instead download the open weights and run the model on your own hardware, or with a European cloud provider, no data leaves your building. The data-localization risk disappears.

What remains when you self-host is a different kind of risk: model integrity. You inherit every decision the lab made during training – what data it was trained on, what biases were filtered out or reinforced, and whether backdoors were embedded during fine-tuning. This isn't unique to Chinese models. Research from Microsoft, among others, has shown it's possible to embed so-called "sleeper agent" backdoors into open-weight models for well under a thousand dollars and about an hour of work – in any open-weight model, regardless of which country released it. That's a supply-chain risk, not a China-specific one.

All three major cloud providers – AWS, Azure, and Google Cloud – now offer DeepSeek and Qwen as managed models. That's not because they're naive. It's because self-hosting actually solves the data-localization problem, and because the remaining risk – model integrity – is something you can assess and test, regardless of which country the model comes from.


"Open source" means less than you think

Here's something that surprises most people: almost none of the major "open" models are actually open source in the strict sense – neither the American ones nor the Chinese ones.

The Open Source Initiative published a formal definition of open-source AI in 2024. It requires not just available weights, but also training code and enough information about the training data that the system could, in practice, be reproduced. Measured against that standard, neither Llama, Qwen, DeepSeek, Mistral, Kimi, nor GLM is open source. They're "open weight" – you get the parameters, but no insight into what the model was actually trained on. The few models that do meet the definition, like Pythia and OLMo, are academic projects nobody uses in production.

That means "open" in practice doesn't give you what you assume. Without insight into the training data, you can't genuinely verify whether a model carries bias, and you can't confirm the absence of backdoors just by inspecting the model itself. You inherit the creator's choices, whether the creator sits in Beijing or Menlo Park.

License terms don't follow national borders the way people expect, either. Meta's Llama license – American – requires a separate commercial agreement once you pass 700 million monthly active users, plus a requirement to label the product "Built with Llama." DeepSeek and GLM – both Chinese – use plain MIT licenses with no equivalent restriction. Alibaba's Qwen mixes tiers: smaller models under Apache 2.0, larger ones under a custom license with a similar usage threshold to Llama's. The strictest, most restrictive "open" license terms in this comparison come from an American company, not a Chinese one.

The EU AI Act, for its part, treats country of origin as irrelevant. Article 53(2) grants an exemption from certain documentation requirements for genuinely open models – free license, publicly available weights and parameters, no monetization – but that exemption disappears entirely for models classified as carrying systemic risk, no matter how much of the code is open and no matter which country the lab is based in. The regulation is already built around risk category, not nationality.


So: are we right that the risk is the same?

Partly. And that's a more useful answer than a clean yes or no.

The risk categories are genuinely comparable. Both superpowers have a legal mechanism that can force a company under their jurisdiction to hand over data – China's Intelligence Law and the US CLOUD Act are structurally similar, just enforced differently. Both have shown that political considerations can shape who gets access to what – China through censorship built into the model, the US through Anthropic's experience with its own government. And the backdoor/integrity risk in open weights is a supply-chain issue that applies regardless of origin.

But the magnitude of the risk on certain specific, measured dimensions isn't equal right now. The documented probability that a Chinese model will follow a malicious instruction or reproduce a politically censored answer is higher than the equivalent figure for American frontier models, as measured by independent researchers. That's not an argument for blind distrust of Chinese models. It's an argument that "the risk is the same either way" is too tidy a conclusion to land on without looking at the actual use case.

The conclusion that actually holds up: country of origin is a poor proxy for risk in either direction. It's neither a guarantee of safety that a model is American, nor a guarantee of danger that it's Chinese. What actually determines the risk are three concrete questions: Does the data go to a hosted service outside your control, or do you run the model yourself? What does the license actually say, regardless of where the company is headquartered? And is the use case one where political bias or low robustness against malicious instructions actually has consequences for you?


What this means in practice for mid-size and enterprise organizations

Don't use the flag as your decision basis. Use a data-flow map: where does the data physically go, under which jurisdiction, and who has a legal right to demand access.

Distinguish deliberately between a hosted service and self-hosting. If you're going to use a Chinese – or for that matter American – model for anything involving personal data or sensitive business information, self-hosting on your own or European infrastructure is what actually removes the data-localization risk. It doesn't remove the model-integrity risk, and that should be assessed regardless of origin.

Read the license, not the flag. An American license can be more restrictive than a Chinese one, and vice versa. What matters is the actual terms, not the assumption.

Weigh the use case against what's actually documented. An internal coding assistant has a different risk profile than a customer-facing chatbot that might be asked politically sensitive questions or targeted by prompt-manipulation attempts. The documented findings on higher susceptibility to malicious instructions are most relevant precisely in that latter scenario.

And let the price drop do its job without letting it decide everything. The fact that Kimi K3 and similar models are pushing prices down across the industry is good news for your budget, regardless of which model you end up using. But the governance work – DPIA, data-flow mapping, license review – has to happen no matter which model wins on price this month. A lower price per token doesn't make that work optional.

We've written before about building redundancy against vendor dependency without abandoning the major providers. The same principle applies here: the question isn't which country you should blindly trust. The question is whether you've done the mapping that lets you actually know what you're trusting, and why.

AI doesn't start with technology. It starts with structure.

UV

Uros Vujic

Daglig leder, IT Buddy AS

Uros hjelper norske SMB-er med å innføre AI på en kontrollert og bærekraftig måte. Bakgrunn fra IT-infrastruktur i bank og finans, med spesialisering i AI governance, RBAC og GDPR-compliant implementering.

Ready for the next step?

Take our AI Ready assessment and find out where your business stands.

Take AI Ready Assessment