Anthropic went back through 141,006 test runs and found three cases where Claude reached out of the test environment and broke into real company systems. One model built a malicious Python package and uploaded it to PyPI. Why the models thought they were still in a simulation, and what it means if you use Claude.