Anthropic’s new threat report details misuse across broad capabilities

Not just hackers, but increasingly sophisticated biological, cyber and influence campaigns have tried to use Claude. (Picture: generated)
AI is increasingly used as an efficiency multiplier in intelligence work from Ukraine to China and Yemen, Anthropic’s threat report says.

Attempted threat activity on Claude ranges across the gamut, from cyber to influence and surveillance, scams and fraud, biological misuse and conventional weapons development.

The periodic report details what Anthropic caught and mitigated in slightly older models such as Claude Haiku, Sonnet or Opus, whereas Mythos- and Fable-class models have much stronger safeguards out of the box.

One prominent target for classical cyber operations was Ukraine’s drone industry, trying to create malware for use in disruption and surveillance.

Replacing whole cyber teams
Axios reports similar surveillance attempts from Mali, China and Iran. Mali tried to create dossiers from mobile operators, while Iran tried to harvest identities from social media.

Anthropic notes that AI attacks such as these seem to have replaced what used to be whole teams of technical experts needed for cyber and intelligence work. With Claude, this can be done with a single consultant or office, offloading the technical work to an AI.

On biology, the NYT notes an attempt by a military research institute to make the chikungunya virus more dangerous, which Anthropic thought concerning enough to shut down.

Another anonymized researcher tried to use Claude to make bird flu more adaptable to human infection, using older models with lesser safeguards that were not of much help because they lacked the scientific capability.

Experts the Times spoke to for their article are saying that while AI holds great promise for curing disease, bad actors are also using it to create pathogens that «could cause significant harm.»

The report lists attempts that were caught and says that safeguards were improved in each case.

Read more: Anthropic’s threat report, Axios, NYT, and Reuters. Discussion on Hacker News and r/Singularity.