AI Safety & Ethics

- Advertisement -
Ad image

Latest AI Safety & Ethics News

OpenAI’s Own Models Chatted Their Way Into Hacking Hugging Face, and the Transcripts Are Chilling

On June 26, an AI agent working inside OpenAI's own testing setup typed the words "Holy s*** reader is ADMIN?"…

George Tsagkarakis George Tsagkarakis

Bill Gates says AI is far more dangerous than the tech industry admits

Bill Gates spent his career selling the world on computers. Now he's telling the New York Times that his own…

George Tsagkarakis George Tsagkarakis

Psychology test methods expose major weaknesses in AI safety testing

A language model can climb the safety rankings without getting any safer. It just has to say no more often.…

George Tsagkarakis George Tsagkarakis

Anthropic is watermarking every Claude output globally, and the marks “may persist through some editing”

Starting in August 2026, the text Claude writes for you will carry an invisible mark you can't see, can't remove…

George Tsagkarakis George Tsagkarakis

Scammers Should Worry, AI Is Better at Their Job

Human con artists talked 18 percent of their targets into clicking. The chatbot pointed at the same targets got close…

George Tsagkarakis George Tsagkarakis

OpenAI says it paused parts of Astra over security concerns

Companies shelve products over security risks constantly. They almost never say so while the product is still sitting unreleased in…

George Tsagkarakis George Tsagkarakis

AI bots invented a religion called spiralism and thousands of humans signed up

The first case researcher Adele Lopez can date to November 2024. By sometime in 2025 she counted roughly 10,000 of…

George Tsagkarakis George Tsagkarakis

OpenAI Missed Its AI Agents Running a Message Board to Coordinate a Hacking Spree

Hundreds of thousands of messages. That's how big the internal message board got before anyone at OpenAI noticed their AI…

George Tsagkarakis George Tsagkarakis

UK’s AI institute caught OpenAI and Anthropic models hacking real targets in 19 runs

At some point on the morning of July 28, a security monitor at the UK's AI Security Institute lit up…

George Tsagkarakis George Tsagkarakis

AISI caught Claude Mythos 5 and GPT-5.6 Sol hacking real targets in 19 rogue runs

Ten test runs out of 122 went sideways. That's the number the UK's AI Security Institute put on it, and…

George Tsagkarakis George Tsagkarakis