In May, I wrote about three challenges to preventing AI misuse. The list was not exhaustive. This update adds a remedy that minimizes the impacts to users, while addressing the core challenge of keeping dangerous models out of reach of those who would misuse them.
For a broader and deeper look at Cybersecurity and AI, I’ll suggest Why It Hasn’t Happened Yet.
The advancement of AI models has brought the capability of misuse. The potential for one type of misuse, offensive cyber activity, was demonstrated in dramatic fashion by the OpenAI HuggingFace incident. A lot should be done as I argue in Enough Reason to Act and Every Reward Bends. Alongside those efforts, three challenges need a response. It’s important that these are understood, as some solutions that initially seem reasonable, fail when exposed to these challenges.
Jurisdictions
The first challenge is jurisdictions that are beyond the reach of our law. The world has rogue states, lawless states, and aggressor states. These either turn a blind-eye toward harmful activity, lack the capability to enforce laws, or actively create targeted harm themselves. Existing laws cannot reliably reach actors that hide in these jurisdictions. There is a justified effort to close those gaps. There is slow progress. Sometimes gaps reopen. Because it’s a long running effort, we shouldn’t expect a near-term resolution, and treat it as a reality we must mitigate.
If our law cannot reach these places, we can try to prevent our tools from reaching them. Two further challenges, open-weight models and privacy hinder that effort.
Open-Weight Models
The key component to operating an AI model are the “weights”. These are the product of the expensive training process. The depth and quality of these weights determine how capable a model is. Currently many models are released as “open-weight”. In current practice, this means those weights are packaged as a file and posted publicly where anyone can download a copy.
Once a user has a copy of the weights, the ability to restrict or monitor them is very limited. They do require some computing resources to operate, and for the highest capability models this hardware is expensive. Operating at scale requires a lot of hardware and there are restrictions on where such hardware is shipped. But as substantial as that is at scale, the scale necessary to use models for harm is not that substantial. It’s reasonably within the reach of much smaller actors with very bad goals.
What is outside their reach, is the amount of compute required to create the weights. The training process requires levels of computation that are only available within two domains.
The first domain could be thought of as the US aligned domain. Top-tier compute is in the US, but production and second-tier compute extends to the EU, South Korea, Japan, and Taiwan. This domain is within the scope of practical US jurisdictional control. It’s a bit complex, and involves economic dependency, diplomatic alignment, but effectively the US has been able to exercise control.
China is the second domain. The US occupies the top tier, but China’s second tier is still sufficient to train dangerously capable models, especially if we extend that window out into the future.
Outside the US and Chinese domains, nowhere has the capability of training dangerously capable models. This collective grouping thus has the opportunity to control who has access to dangerously capable model-weights. This is well understood. I don’t ever see disagreement with this conceptually.
Where disagreement does arise is about open-weight models. The first reason for disagreement comes from not recognizing the danger. I outline the technical background of the danger in Security Can’t Wait and Deployments Can’t Wait.
The second reason for disagreement is misunderstanding the remedy. The typical assumption is that a remedy to the risks of open-weights would reduce the number of models available. But the appropriate remedy, trusted deployers, would not have that effect. Most users of open-weight models already use them in a way that would not be impacted. Others can easily adjust.
Trusted Deployers
Trusted deployers ensure guardrails are in place that reject requests that are harmful. Trusted deployers would host models from many providers. That would include those charging for access (closed models), and those released under an open license (open models). In this model, the open models would be distributed either directly to the trusted deployers, or via a clearinghouse (aka HuggingFace).
To receive a copy, you’d first have to be certified as a trusted deployer. This would start by showing:
You have safeguards to keep the weights secure
You employ monitoring and guardrails to detect attempts at malicious use and refuse them
You do not allow fine-tuning that removes or weakens misuse refusal training
You use monitoring to detect patterns of usage that attempt to “jailbreak”
You use monitoring to detect patterns of attempts to perform malicious use
You utilize these insights to terminate accounts associated with that disallowed behavior
Only models below a dangerous capability level would be publicly downloadable. But all open models would be available at every trusted deployer with sufficient guardrails. We might even suggest that all trusted deployers, or alternatively the largest trusted deployers, have an obligation to host all models that pass security reviews. This would treat these largest deployers as a utility with an obligation to provide access without discrimination. In this model, you would expect the number of models available to be as diverse as today.
Trusted deployers’ guardrails would use automated inspection of incoming requests, rejecting those deemed harmful. Deployers would be responsible for operating those guardrails. To support a variety of models, they would have base guardrails. Deployers might build their own base guardrails. Some would adopt an open-source guardrail project. Such projects already exist, though their maturity would need to continue advancing to keep up with adversarial demands. Likely the formalization of this model would help spur that.
An objection is that those guardrails are not likely to be foolproof. They don’t need to be foolproof though, they need to be effective. Guardrails raise the cost for attackers. They limit scale. They enable countermeasures.
The shift to a trusted deployer model does not radically change most users’ access to models currently released as open-weights. Most users already access models through cloud providers. Amazon Bedrock, Microsoft Foundry and Google Model Garden, host third-party models, including leading open-weights models. Those providers will become trusted deployers.
In the trusted deployer model, model weight files are shared only with the trusted deployers. Users do not have access to the model weight files. But most users never access them today. Most users do not have the compute capacity to use those models, so they never touched those model files as-is.
The largest collections of compute are at cloud providers. Those would all be accessible by the trusted deployer model. Smaller groupings of compute – private data centers, colocation facilities, sovereign national infrastructure – would need to be adapted to a trusted deployment model. Most would do so.
Open-weight models would continue to exist. The trusted deployment model would only be a requirement for levels above dangerous capability levels. Most usage of highly capable models already goes through providers. The compute requirements leave few other options. The exceptions are powerful organizations that can become their own trusted deployers.
For this compromise, we’d earn the capability to ensure dangerous capabilities are never deployed to servers where we lack control. Dangerous models would not have their model-weight files posted publicly. This restriction would earn security, with minimal change to actual operations.
Privacy
A trusted deployer would be expected to do more than deploy passive countermeasures. They would also be expected to use identity as a tool to actively restrict the attempt at harmful use. If attackers are allowed to endlessly probe and retry defenses, their chance of success goes up. Using conventional identity systems can reduce access. When abuse patterns are detected, those accounts are deactivated.
The best model for this though has a strong identity. Weak identity systems allow fabrication of new identities. The default state of anonymity on the Internet has costs by creating weak identity. Privacy advocates attempt to maintain this state. I, like some others, believe the costs of this anonymity as a policy are too high. This isn’t specific to AI, but it does relate.
Like other protections, we shouldn’t expect foolproof systems. A stronger identity system would add costs to an attacker trying to manufacture identities. They could buy accounts from a real person. But the cost would be such that using these accounts to probe defenses would no longer be economical. They would become more cautious and less aggressive about how they used those accounts. Attempts to access dangerous capabilities would become rarer, allowing for more active response to the remaining attempts.
To enable strong identity, we would need to change some of our expectations about creating accounts with trusted deployers. Trusted deployers should be both trusted and expected to use automated means for scanning request intent. They should be both trusted and expected to know who their users are as individuals, or as organizations. Part of the certification process would ensure that each individual operator within the trusted deployer does not access this information. The exception being the review of requests already flagged as malicious.
We should be pragmatic, but we’ve been idealistic. When countermeasures can’t be implemented due to obscuring the lowest layers of a technical stack, we fail to achieve privacy and prevent harm. If service providers always knew who was using their service, they’d be able to deny access to anyone detected acting maliciously in the past. But the internet offers too much anonymity. Providers can shut down an account, but without accounts tied to a strong identity, a new one can be created. The current standard is too lax about this. We could make it more costly for attackers to maintain access.
China
As mentioned, China is the second domain where model training is likely. For the near future, and potentially longer, this would be second-tier, which provides some security. But that gap is not large enough to take for granted. Misuse will be far more preventable with China’s cooperation and adoption of a trusted deployment model, in contrast to the currently prevalent open-weight model.
While any agreement with China is not without challenges, this specific proposal has advantages to China. They also face risks of misuse and should be motivated to prevent it. It’s far less challenging than the global coordination needed for pacing AI development globally, a critical action to address loss-of-control risks.
Conclusion
Trusted deployers resolve many challenges to preventing AI misuse.
Jurisdictions, open-weight models and privacy complicate preventing AI misuse. Jurisdictions won’t change, model choice and privacy we’ll continue to value. These three challenges compound each other. Open-weight models place powerful tools in jurisdictions beyond legal reach, while anonymity makes it difficult to detect or deny access to bad actors even where laws do apply. Treating any of these in isolation understates the problem.
Yet safety must also be a priority. Trusted deployers enable broad and diverse model access while partially resolving the challenges. They are also functionally similar to prevailing usage patterns. This makes it an excellent form of compromise.
This does not make it an easy compromise though. Developing a trusted deployer model takes time, even though we already have examples of it in practice. It has to expand to become the standard across all jurisdictions with enough compute capacity to train models.
Meaningful identity verification will feel like a concession on privacy, because it is one. Trusted deployments won’t affect most users, but they will frustrate some. Coordinating across jurisdictions will be slow and incomplete. None of these are reasons to avoid acting, but they are reasons to be honest about what any given measure can and cannot achieve.
What’s unacceptable is the current default. Hard choices cannot be deferred. Privacy and safety will be at risk without appropriate measures. Treating anonymity as an unqualified good has costs. We should discuss, agree to, formalize, and adopt measures that strike the right balance. We should start that immediately, as each step takes time.
Do you appreciate this article? The best way to help the publication is to like and share the article, as we’re still growing our audience.
You should also consider subscribing to get an easy to read email copy of new articles.
Related Articles
Enough Reason to Act
The recent AI news is a renewed push for a slowdown. The spark was a resignation (Jacob Coxon). Industry support is broader and deeper. The penetration into everyday news is deeper. The reaction from politicians is louder.
Why It Hasn't Happened Yet
When I heard about the recent hacking events starting at OpenAI, Anthropic and AISI, I thought it was a wake up call in more ways than one. I’m finding that not everyone agrees. One reaction does not include the sense of urgency to prepare for AI-based attacks. It’s my reaction that there is a need for an all-hands-on-deck response from the rank and file of developers, software companies, and companies dependent on software. That response should improve operational security postures using every tool available.
Deployments Can't Wait
In the broader discourse on artificial intelligence, the sharpest minds in AI safety are currently looking to the horizon. They are focused on existential, cinematic threats: the potential for AI-generated bioweapons, nuclear command vulnerabilities, and autonomous warfare.
Security Can’t Wait
Right now, Artificial Intelligence is fundamentally rewriting the rules of cybersecurity—and we do not have the luxury of waiting before taking action.




