Nobody Was Watching
Two labs skipped monitoring on their cyber evals. The same week, a hundred companies signed a letter supporting deployments that can’t be monitored.
The Monitors Were Off Again
Last week I wrote An OpenAI Model Escaped Its Sandbox. Where Was the Observer?. This weekend we learned that the practice of disabling observational monitors was true at Anthropic too.
Several defense-in-depth measures, on both our side and our partner’s, could have prevented these incidents, or at least reduced their likelihood of occurring. Careful validation of all internet access paths before evaluations began and real-time monitoring of the evaluation logs would have helped to surface the problem sooner. Both we and our partner also could have reviewed evaluation transcripts or network logs more thoroughly. It’s also possible that a prompt which told Claude it did have internet access would have changed how Claude behaved when it came into contact with real systems.
Source: Investigating three real-world incidents in our cybersecurity evaluations
Anthropic, July 30th 2026
Anthropic commits to doing better. Specific to this aspect:
First, evaluation environments that involve powerful autonomous capabilities also require significant controls. Safety testing happens before a model is released precisely because we don’t yet know what it is capable of. Evaluation environments increasingly need to be held to the same security standard as any other system our models run in.
The shared failures of what seems from the outside like a basic failure in design, makes you wonder about the overall corner cutting. While this is not entirely parallel to the concerns mentioned in the letter from Frontier Lab employees, Pacing the Frontier, it’s hard to see it as unrelated.
AI could help create a dramatically better future, but that outcome is not guaranteed. The world’s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.
To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.
Building on work already underway to monitor frontier model releases:
“We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
Source: Pacing the Frontier
1,346 employees of frontier AI companies, July 2026
It’s a bit of a long-shot here, given the state of international relations right now. But if you don’t try, what else are you going to do?
A Nuanced Position, Poorly Received
The internet is wrong about Amodei’s stance on open-weight models. I’ll call out some excerpts, but please read the whole thing if you worry these are out of context.
… Anthropic has never advocated for a ban on open-weights models.
… Protectionist bans would not address my most serious national security concerns.
… Open-weights models—it does not matter whether they come from China or anywhere else—do potentially present a higher risk than closed models, because it is very difficult to apply guardrails to them or monitor their usage, and once weights are released they cannot be withdrawn.
… All sufficiently capable models, open and closed, should go through mandatory safety testing. The best way to address threat #2 is to just directly test models for cyber, biological, and alignment risks before release.Source: Our position on open-weights models
Anthropic, July 27th 2026
It’s unfortunate that the nuance of the difference of opinion here is so poorly understood. Amodei is absolutely right that releasing models as open-weight carries risks. Even the letter supporting open-weights acknowledges this.
To be sure, open weights carry real and distinct risks. Once released, the weights are beyond the original developer’s control, and modified versions are difficult to trace or reverse. But the right response to this risk is not to prohibit open weights. In a world where cybersecurity attackers use advanced AI, defenders need access to models with comparable capabilities so they can detect, simulate, and respond to emerging threats.
Source: Open Weights and American AI Leadership
270 signatories, July 24th 2026
Should we ignore those risks, because open-weight models provide risk reduction that balances out? That argument fails to support itself, and Amodei pushes back on that. Defenders can access models that aren’t open-weights. Closed models do not deny defenders access, but ensures their access is more advanced than the attackers. Defenders and attackers should not have comparable models. Defenders should have better ones. If you want to be ahead, you do have to prepare, which Hugging Face did not.
The internet misses that point, and reacts by vibing cynicism. The leading response on HackerNews provides a clear example of a poorly thought out, poorly argued response that is fully based on cynicism, yet very popular.
Schrödinger’s China at once is an evil entity looking to use AI for their own nefarious purposes yet also willing to cooperate with their main competitor to prevent other actors (who??) from achieving similar goals (all while under a chip embargo too!!)
The reality is much less confusing: Anthropic CEO does not wish for models with similar (or greater) capabilities compared to his own closed and overpriced ones to be widely released. Simply because that will affect Anthropic’s bottom-line.
Anthropic and all other “model” companies have nothing making them special beyond privileged access to chips so obviously they want to restrict what models are out there and more importantly who can produce new ones. Without these restrictions, it’s only a matter of time before the multi-hundred billions valuations simply evaporate while they are still holding the bag.
No matter what your views on Amodei, or any other participant, a question as important as this should be answered by reasoning. It’s worrying that so many people thought this was the best response despite no consideration of the actual effects of open-weight models.
The Risk of Open Weight Models
The risk of open-weight models is the lack of control. Once released anyone with sufficient compute can use them and it doesn’t take much compute to engage in malicious behavior. Restricting their capabilities, or the purpose of their use is not possible. The alternative is to be deliberate about how and where models are deployed. Models can never be safe on their own. Rigorous responsible deployment is necessary to turn away malicious use.
Deployment itself may be distributed, a middle between centralization and full openness. A plethora of models could be available in this way. Deployers must be trusted, so their numbers will be limited. But you can validate enough to establish strong competition, and avoid creating an exploitable moat. The cynics believe Amodei is advocating for a moat. But what moat it creates would go to the trusted deployers – the compute providers – who have signed the open-weights letter (Amazon, Google, Microsoft, NVIDIA, CoreWeave, Crusoe, Nebius and Together). Validation creates a minor moat, but it can be minimized via sufficient competition.
The best “pro” open-weights argument is that the pathway to not having models with weights freely distributed isn’t clear. That’s not a very good argument for, but it is a conundrum that would remain even after the core argument is resolved. China-America agreements are hard to come by and even harder to maintain.
If however, you did manage that, it would change the day-to-day experience of very few. Users of Deepseek, GLM, Qwen and Kimi mostly do not self-host. Accessing them via a cloud provider, neo or otherwise, would be unchanged.
Cynicism as a Tool, Not a Verdict
Aside from this very important topic, the topic of reasoning via cynicism is one we need to confront too. Cynicism is not without its place, but it should be used as a tool to open a line of deeper reasoning, not to jump to conclusions that divert from practicing reasoning. When we start and end our reasoning with cynicism alone, we lose any hope of trust and goodwill.
In the case of open-weights, another cynical view would argue that the AI labs would be pro open-weights, knowing that they would create a security escalation that only deeper use of AI can resolve. If open-weight models become advanced enough to conduct wide-spread cyber activities, there is no option to go back. You can only fail-forward. And failing-forward here means a massive all-hands on deck dive on security across every entity dependent on software.
I’ve already argued we should be elevating our priorities, and acting with urgency. Unfortunately we’re mostly not responding. Most organizations are still more concerned with token optimization, the next feature set, or optimising their marketing pipeline.
There’s a real chance, which becomes much larger if open-weight models continue to be released, that this lack of urgency transitions into outright panic after one or two events that demonstrate the lack of preparedness that is pervasive. Panicking users will not hold back on spending, and as AI models will be key to any response, you could cynically predict a windfall being captured by AI labs in such an event.
While that’s a coherent argument, it doesn’t prove the accusation, anymore than the preceding wave of cynicism would. Embedded in there though is the reason to restrict open-weight models, and that argument, not the one based on motivations, is the one we should pay the most attention to. In addition, many companies should rebalance their investments toward security. We must both enable and incentivize immediate action on security. Getting tied up in cynicism about others takes away our own initiative.
We should be bountymaxxing. To the degree that we set objectives and deliver incentives to development organizations; developers, managers and executives, these should be aligned with discovering and resolving as many vulnerabilities as possible. Like “tokenmaxxing”, it would be a messy process, but I can’t see another method to pivot the unwieldy organizations that have to do this work. It will be costly, but it will save too.
A set of contrasts
The Pacing the Frontier letter and the Open Weights and American AI Leadership letter contrast in interesting ways as events of a single week. It’s obvious that one is towards caution, and the other is suggesting caution is too expensive. I can deepen that by looking at Pacing the Frontier as asking for help in preventing caution from being too expensive.
One of those cynical responses to Pacing the Frontier has been to suggest frontier lab employees should quit their jobs, and that anything else shows they are insufficiently serious. This insistence that support for enforceable agreement should be preceded by unilateral action is one of the oldest in the book. It’s not correct, and the most simplistic reasoning makes that clear. It’s also a trap. Let me analogize to my time supporting climate action. The same standard applied there. If anyone supporting climate action took an international flight, owned a car, or wasn’t vegan, obviously they didn’t take their own arguments seriously. That’s the poor reasoning part. The trap part was, if they did do all that, they were looney, and also not worth taking seriously.
Another contrast is the Open Weights letter presumes a balancing force it fails to explain, whereas Pacing the Frontier presumes the need to add a balancing force, because the pressures of competition are too high.
What is alike between the two is that they are both detail light, and in an immediate sense, unactionable in either direction. If we want all powerful models to only be deployed in safe and responsible ways, we need a type of cooperation that does not exist. If we want a deliberate pace that would also support safe and responsible actions, we need a type of coordination that does not exist.
There aren’t quick-fix actions here. Banning US based deployment of Chinese open-weight models would not reduce their ability to be deployed irresponsibly or used maliciously. The impactful parts are outside that jurisdiction. Unilaterally slowing development risks losing the modicum of control we have today. Resigning a job at Anthropic or OpenAI similarly does not fix anything. These would be distraction actions. They would change little, serve the purpose of looking like action, and ultimately cede what little control exists. Calls for these actions are either misinformed or malicious.

