AI

The AI Models are Open… To Anything

Chart comparing open-weight AI model capabilities versus safety refusal rates in 2026
Chart comparing open-weight AI model capabilities versus safety refusal rates in 2026

There’s a working assumption in most AI conversations that the most powerful models come from a small group of well-resourced labs, they’re locked behind APIs, and those labs have at least some incentive to make sure their systems don’t help people build bioweapons. That assumption is getting harder to defend.

A new evaluation from SaferAI, cited in a TechCrunch report published August 4, found that GLM-5.2, an open-weight model from Chinese startup Z.ai, is only months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 on the capability metrics that matter most for doing harm. We’re talking specifically about offensive cybersecurity tasks and biological threat knowledge. GLM-5.2 refused none of the tasks in SaferAI’s testing. Claude Opus 4.7, for reference, refused so consistently that SaferAI couldn’t even complete their CyberGym evaluation on it.

That’s not the gap most people expected to see.

The Capability Gap Is Closing Faster Than Expected

On the standard benchmarks developers and researchers use to rank models, the convergence between open-weight and closed frontier models has been significant. DeepSeek V4-Pro, which anyone can download and run locally under an MIT license, scores 82.6% on SWE-Bench Verified, a coding benchmark. GPT-5.5 sits at 88.7%. That’s a six-point gap on a test that two years ago looked like an uncrossable moat between the big labs and everyone else.

More broadly, open-weight model quality has closed to within 5 to 15 points of the closed frontier across major benchmarks. Open-weight models do this at a fraction of the per-token cost of closed APIs. Even if you accept that the closed models are still meaningfully better at the highest levels of complex reasoning, the gap is now measured in months, not years.

I’ve written before about how benchmark results aren’t always what they seem and how the process of running those evaluations isn’t always clean. That context matters here. When we say open-weight models are “closing the gap,” we should be careful about which gap we mean. But the SaferAI evaluation isn’t about leaderboard bragging rights. It’s asking whether the models capable of helping with genuinely dangerous tasks are also the ones least likely to refuse. Based on available evidence, the answer is yes.

Benchmark comparison showing open-weight models closing gap with closed frontier AI models

Safety Is Running on a Separate Track

Here’s where things get uncomfortable. A research group called Far.ai identified hundreds of universal jailbreaks that work across frontier closed models, including xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. So the safety controls at the big labs aren’t perfect either. Jailbreaks using roleplaying, authority impersonation, fake conversation histories, and follow-up prompt chains continue to work at meaningful rates.

But there’s a structural difference between a closed model and an open-weight one. When a jailbreak works on GPT-5.5, OpenAI can see it happening at the API level, patch it, and push an update. When a jailbreak works on a locally hosted open-weight model, there’s no one on the other end of the connection. The model just runs. Whatever safety fine-tuning was applied during training is the only protection that exists, and that protection can be stripped out entirely by whoever downloaded the weights and wants to fine-tune the model further.

Z.ai hasn’t published a safety framework, pre-deployment testing commitments, or a risk assessment for GLM-5.2. That’s not a minor oversight when the model in question performs comparably to the best closed models on tasks specifically designed to test capability for harm.

Henry Papadatos, executive director at SaferAI, put it plainly: “The frontier of capability is not the frontier of risk.”

That sentence is worth sitting with.

Who’s Responsible When Nobody’s In Charge

The open-source AI community has legitimate arguments for why open-weight models matter. Clem Delangue, CEO of Hugging Face, has pointed to their value for cybersecurity defense research, among other uses. The ability to run a model locally, inspect its weights, and customize its behavior is genuinely useful for researchers, developers, and organizations with legitimate data privacy needs.

None of that is wrong. But it doesn’t resolve the question of who’s responsible when an open-weight model causes harm. Right now, the answer is nobody in particular.

Graham Webster from Stanford’s Cyber Policy Center noted that Chinese regulations around AI tend to focus on political content rather than catastrophic risk potential. That’s not unique to China. Most regulatory frameworks are still fighting the last war: focused on bias and misinformation while the actual frontier concern is whether a model can help someone who doesn’t already know how to synthesize a dangerous pathogen figure it out faster.

The regulatory arbitrage here is real. If a lab’s weights are publicly downloadable and the lab faces no meaningful accountability for downstream use, it has little incentive to build in safety controls that might limit the model’s appeal. The gap between what a model can do and what it’s designed to refuse isn’t a technical problem at this point. It’s a governance problem, and nobody’s governing it.

What This Actually Means Going Forward

Shutting down open-weight AI development isn’t the right answer. The research benefits are real, and the access it provides to smaller teams and academic institutions actually matters for the long-term health of the field. But “open development is good for innovation” and “releasing a capable model with no safety framework is fine” are two different claims, and we’ve been treating them as the same claim.

The question nobody wants to answer is what accountability looks like for a model that’s been downloaded millions of times and is running on hardware nobody controls. Closed API providers can at least be held responsible at the point of access. With weights distributed publicly, that point of control is gone.

We’re going to need a better answer than “the community will sort it out.” The models are getting better every quarter, and the community hasn’t sorted it out yet.

Sources

Follow Chris Meredith on Medium at @ChristopherMeredith

Chris Meredith writes about AI, technology, and what it actually means for real people. Follow along on Substack: monkeyattack.substack.com

Leave a Reply

Your email address will not be published. Required fields are marked *