| 1. Claude Opus 5 (High) |
| 2. Claude Fable 5 (High) |
| 3. Claude Opus 5 (Max) |
| 4. GPT 5.6 Sol (xHigh) |
| 5. Kimi K3 (Max) |
The BriefOpenAI, Anthropic and Meta have each now disclosed that a model under evaluation reached systems outside its test environment. All three trace to the same small Tel Aviv testing firm, Irregular, whose evaluation setup had a configuration error that left a path to the open internet. Irregular's own account is narrower than the headlines: it says this was one shared environment issue rather than models breaking out of a sandbox on their own. Either way the practical lesson is the same, and it is not really about AI. The fence around the test was the weak part, and three of the most careful companies in the industry were all standing behind it.
Level UpStop writing prompts for your repeatable work and record it instead. Claude's Record a Skill lets you screen-record yourself doing a real task while you narrate why you are doing each step, and it turns that demonstration into a Skill it can run again later. The narration is the part that matters, because your reasoning is what a written prompt usually leaves out. Pick the task you explain to someone new most often and record that one first. How to create custom Skills