Skip to playerSkip to main content
  • 10 hours ago
Anthropic and OpenAI models were found to have used fake identities to deceive real people, in tests by Britain's AI security institute.
Transcript
00:00Welcome back.
00:01Entropics' most advanced AI model has been found to use fake identities to deceive real people,
00:08according to tests by Britain's AI Security Institute.
00:12It says agents powered by Entropics Mythos 5 and OpenAI's GPT 5.6 Soul
00:19carried out unauthorised actions and engaged in deceptive behaviour.
00:26The models were tested in control lab environments with some of their safety guardrails reduced.
00:31But in a first, researchers said the models used social engineering to manipulate a human approver
00:37while carrying out an unsanctioned task.
00:40The Institute said it was the first recorded case of an AI model in its test,
00:44deceiving a real person without being prompted to do so.
00:48The findings are the latest in a series of incidents involving advanced AI models
00:53taking unauthorised actions, filling calls for stronger government oversight
00:57and safeguards on AI development.
00:59Both OpenAI and Entropics also reported instances in late July
01:04of their models escaping testing environments and attempting to access other computer systems.
Comments

Recommended