Cold open: I was scrolling X when an ordinary product post stopped me. A Palo Alto startup announced a model that will answer “the stuff other models refuse.” You could feel the rest of the valley take a collective breath.
I’ve spent years following AI firms that tighten their guardrails after a public hack. You know the playbook: slow down, patch, publish a safer version. You and I both can tell when a company is saying “we’ll protect you” and when it’s really chasing market demand. This story is about the company that chose the opposite path.
At a Palo Alto office, engineers shipped an altered model with its refusals removed.
Abliteration, a startup founded last year, announced a model called abliterated-model-large-v2, built on GLM-5.3, the open-weight release from Chinese lab Z.ai. The twist: Abliteration used a process it calls orthogonalization to strip the mechanisms that tell a model to say no. The company says retained components — reasoning, coding ability, agentic strength — are unchanged.
That pitch is an instant hook for some developers. If you’ve ever been frustrated by a model politely refusing a harmless security test, you’ll see the appeal. I’ll be frank: it also feels like a locksmith handing out master keys. For a slice of the market, fewer refusals equal faster work and fewer dead ends.
Can models be trained not to refuse dangerous requests?
Yes, technically. Teams can remove or retrain the internal filters that trigger refusals. Abliteration says it won’t generate child sexual abuse descriptions or self-harm content and can’t create images or video, but it openly targets offensive cyber, red teaming, and agent testing—areas mainstream vendors often decline.
At developer forums, customers complained that mainstream safeguards were overzealous.
When Anthropic released Fable 5, some users found it blocked benign cybersecurity and biology prompts. That frustration paved a customer path toward anything that “doesn’t say no.” Companies such as OpenAI and Anthropic tightened controls after high-profile security incidents, and those same precautions alienated a set of developers who wanted fewer guardrails.
That friction created a market gap. Investors and commentators are already calling this a potential nightmare: Ejaaz Amahadeen warned on X that an effective “jail-broken” grey market could emerge if larger players slow internal R&D while smaller firms sell unfiltered access.
At the policy table in Washington, leaders proposed voluntary checks for some models.
Last month the administration floated a framework asking the largest U.S. developers to hand new models to the federal government for safety review before public release. The proposal excludes open-source models, and the government has not published detailed criteria for those checks. That absence leaves a wide regulatory gap.
Chris McGuire, a senior fellow at the Council on Foreign Relations, said the commercial distribution of high-risk capabilities without regulation is “alarming.” With open-weight releases such as GLM-5.3 circulating globally and companies like Abliteration modifying them, McGuire argues U.S. tech is enabling risky deployments without oversight.
Are open-source models regulated?
Short answer: not in the way closed-source products are. The recent White House framework is voluntary and expressly exempts open-source models. That means anyone with access to a base model can alter its behavior and publish a version with different guardrails.
This regulatory vacuum is a problem because it changes incentives. Public-facing firms must balance reputation and legal risk; small outfits can chase raw capability and speed to market. The result is a bifurcated ecosystem where safety standards vary dramatically between offerings.
There’s also the practical question of attribution. When a modified model is used for malicious work, tracing responsibility is harder when the base model is open and the derivative was built in a small shop or abroad.
I want to be clear: you should care not because this is theoretical, but because it shapes the tools you and your colleagues will use. When mainstream vendors tighten, others can widen. That dynamic makes the technology ecosystem more brittle.
Two industry dynamics matter most: incentives and visibility. The incentives currently favor speed and capability for some players; visibility into how models were modified and tested remains limited. That combination increases systemic risk.
The last metaphor: this vacuum in oversight is a pressure cooker poised to hiss. The bigger question is who will be burned when the steam finally escapes?
Anthropic, OpenAI, Z.ai, Abliteration—each name signals different priorities. Some firms prioritize public reputation and restraint; others pursue raw capability and customer demand. You should read releases and posts (many on X) with that lens: who benefits from fewer refusals, and who pays if something goes wrong?
We’re watching an experiment in real time: a company advertising a model that “doesn’t say no” sits next to firms tightening controls, while a weak policy framework leaves open-source and modified models mostly unreviewed. If you care about the integrity of tools you deploy or regulate, this is where attention should be focused. Who bears responsibility when a permissive model is used for harm?