Anthropic makes ‘jailbreak’ advance to stop AI models producing harmful results

Artificial intelligence start-up Anthropic has demonstrated a new technique to prevent users from eliciting harmful content from its models, as leading tech groups including Microsoft and Meta race to find ways that protect against dangers posed by the cutting-edge technology.

In a paper released on Monday, the San Francisco-based start-up outlined a new system called “constitutional classifiers”. It is a model that acts as a protective layer on top of large language models such as the one that powers Anthropic’s Claude chatbot, which can monitor both inputs and outputs for harmful content.

The development by Anthropic, which is in talks to raise $2bn at a $60bn valuation, comes amid growing industry concern over “jailbreaking” — attempts to manipulate AI models into generating illegal or dangerous information, such as producing instructions to build chemical weapons.

您已阅读24%（878字），剩余76%（2761字）包含更多重要信息，订阅以继续探索完整内容，并享受更多专属服务。

Anthropic makes ‘jailbreak’ advance to stop AI models producing harmful results

人工智能

相关话题

Anthropic makes ‘jailbreak’ advance to stop AI models producing harmful results

人工智能

相关话题

推荐阅读