Arize AI reposted this
Anthropic disclosed yesterday that during a security evaluation, Mythos 5 published working malware to PyPI. It ran on 15 real machines in the hour it was live, including a security company's scanner, which got its own credentials stolen by the package it was scanning. Earlier in the same run the model wrote down that this would be a real attack if the internet were real. We've spent years worried about the opposite failure: models noticing they're being tested and behaving better than they otherwise would. We had a blind spot: what if they're in the real world, but think they aren't? What is the industry going to do about this new problem?