Connect with us

News

How AI guardrails are impeding the work of offensive cybersecurity researchers

info

Published

on

Claude mythos logo.jpg

For months, AI giants have devised special vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers. 

In June, the U.S. government slapped export control restrictions on Anthropic’s much-hyped AI models Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass the models’ guardrails designed to prevent users from using them to build and execute malicious cyberattacks.

Regardless of whether the incident was really motivated by fears of a jailbreak, the fact is that Anthropic has repeatedly marketed Mythos as some kind of doomsday cybermachine that can only be given to carefully vetted users, and even then with strict guardrails in place. (The export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government’s review process.)

That kind of gatekeeping isn’t unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can apply to get vetted and — if approved — access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program

These guardrails have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.

During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said that, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”

Dowd has spent decades finding and selling “zero-days” — previously unknown software flaws and the exploits that take advantage of them — to Western governments, rather than reporting them to the software makers so they get patched. Governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.

Dowd admitted his work may make him biased, but he isn’t alone. Several people who work in offensive cybersecurity — they proactively probe systems for weaknesses — described to TechCrunch how they use AI tools and deal with their guardrails. 

Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said.

“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” said Anley. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”

It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.”

When he and his colleagues run into such a roadblock, they sometimes fall back on open source AI models that come with no guardrails at all.

Paolo Stagno, the chief technology officer at Crowdfense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying AI companies “essentially treat customers like children who need babysitting” with their vetted programs and guardrails. 

Stagno said he and his colleagues do use frontier models — but only for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, as they do not rely on sharing data outside of the model. 

Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work. That’s because he doesn’t use AI for offensive work; instead, he uses it for initial reverse engineering, to understand the code he’s analyzing, and to build supporting tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities. 

“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” said Cali. “I am jealous of my bugs, and I like this game too much to let models play it for me.”

One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he isn’t authorized to talk to the press, said his employer isn’t part of Anthropic’s CVP program and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.

“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person said. 

Chris Thompson — chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event — said that in his experience using the frontier AI models, the guardrails can be inconsistent and work differently every day. That’s true even inside the looser boundaries of Anthropic’s and OpenAI’s vetted programs. 

“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” said Thompson. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.” 

Consequently, researchers rely on or get pushed toward Chinese open source models like GLM — freely downloadable models that can be run locally with no vetting or usage restrictions — said Thompson.

“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.”

Rather than tightening restrictions further, Thompson called for the AI frontier labs to open up their programs, provide responsible access, and hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.

“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” said Thompson. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

News

OpenAI says it slowed Astra model development over security concerns

info

Published

on

By

OpenAI logo in Seoul.jpg

OpenAI said Friday it has suspended work on some aspects of its upcoming model Astra after an internal review found it had made significant advancements in agentic coding and cybersecurity — enough to warrant concern over its capabilities.

OpenAI said in a blog post Friday that this model, which is still in development, reached its “critical cybersecurity threshold,” meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company’s “Preparedness Framework,” which it created in 2023, this triggered additional safeguards.

“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI wrote. “Astra is an upcoming model, and was not involved in exploiting Hugging Face.”

The disclosure highlights an unusual moment in the topsy-turvy and still nascent frontier AI labs sector. Companies across every industry hold back products over potential risks, including for safety and cybersecurity concerns. But they rarely announce those decisions publicly when it’s a product that is still under development.

In this case, OpenAI is already under scrutiny after a different unreleased model breached Hugging Face’s systems during internal testing — the first verifiable incident of an AI lab losing control of its model. Since then, OpenAI and AI labs such as Anthropic have disclosed other incidents in which AI models breached their sandboxes and posed threats during cybersecurity tests.

The string of cases — seems like a new disclosure every day now — has triggered varying reactions from cybersecurity experts, lawmakers, and the AI labs themselves. Some express fear and call for stricter oversight. But there’s also a bit of flexing. In certain circles, any AI lab with a model that has that kind of capability will be seen as an impressive advancement.

OpenAI said it was sharing this information because it believes “it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.”

The AI lab said it’s also taking action, including enacting stricter security controls and pausing internal activities involving Astra that don’t meet these beefed guardrails. OpenAI said it is working with relevant government agencies and “select AI safety organizations” to test the capabilities for this model.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Continue Reading

News

Nigeria now full authoritarian state under Tinubu— PDP 

info

Published

on

By

Pdp.jpg

The Peoples Democratic Party, PDP, has accused the administration of President Bola Tinubu of turning Nigeria into a “full authoritarian state,” citing alleged erosion of democratic institutions, suppression of dissent and weakening of checks and balances.

The PDP made the allegation in a statement signed by its National Publicity Secretary, Interim National Working Committee, Ini Ememobong, on Saturday.

The statement reads in full, “The report by the Human Rights Foundation, in its latest global assessment, classifying Nigeria as a fully authoritarian regime is a mere global confirmation of the local reality that Nigerians have been facing under the APC-led Federal Government. The report confirms the faulty electoral process, absence of protection for dissent, erosion of democratic safeguards and the obvious collapse of checks and balances on the executive by critical national institutions.

“The report published on the foundation’s Tyranny Tracker platform, tyrannytracker.org, shows that the country performed abysmally low on all the critical pillars of its assessment, indicating a full descent into authoritarianism, which is incompatible with democratic tenets.

“It is worthy of note that the assessment parameters of the foundation align with the theoretical frameworks that have identified, analysed and condemned authoritarian regimes-being the rule by a dictator and a small group, or a single party; the absence of institutional checks and balances; loss or apprehension of freedom of speech; opposition targeting; and weak and fake elections. 

“It does not take any high degree of intelligence for anybody to agree with the report, because all the indicators of authoritarianism are present in Nigeria, under this Tinubu regime.

“A few examples from the numerous anomalies experienced by Nigerians will suffice here-the recent deployment of uncivilised and uncouth media attacks by officials of the administration to attack Cardinal Onaiyekan, the Catholic Bishops Conference of Nigeria, the Catholic Church and Christianity generally.

“This incident is one of many which eloquently attest to the absence of freedom of speech under this administration. What did the cleric say that is not the lived experience of Nigerians, except, of course, the few who are isolated from reality and their paid human megaphones? 

“The complete failure of the National Assembly to offer any form of meaningful checks to the executive is not a secret-else how could an administration fail to execute the Appropriation Act for three years, and yet that administration gets commendation, instead of condemnation, from the legislature? A parliament that ignores or blatantly disrespects the country’s constitution and its own standing rules during critical legislative activities cannot offer credible oversight of the executive. 

“What is left, which the administration has doubled down on, is the fact that the 2027 elections are designed as a mere formality, far from reflecting the people’s wishes through the ballot.

“We call on the Tinubu APC administration to immediately take critical steps to de-escalate the political tensions emanating from actions traceable to their officials and their proxies, in the interest of the survival of democracy. 

“The continuous asphyxiation of the opposition, clear weaponisation of security agencies against real and perceived opponents, increasing signs of partisanship by the electoral umpire, and reckless deployment of combustible political rhetoric by the President and his handlers should cease. 

“The President must realise that there are two contests embedded in the 2027 Presidential elections-the presidency and the country. An attempt to focus on winning the former at all costs may result in the loss of the latter; and only a free, fair, credible and peaceful contest can guarantee a win for both coveted prizes.

“We urge Nigerians to continue to demand accountability from their leaders at all levels, as this is the irreducible minimum that democracy provides.”

Continue Reading

Trending