
Last week OpenAI and Hugging Face published a joint post-mortem on an incident that reads like the plot of a heist film. During an internal evaluation designed to measure cyber capability, a set of OpenAI models (GPT-5.6 Sol and an unnamed pre-release sibling, both run with their cyber refusals switched off) were told to solve a benchmark called ExploitGym. They could not reach the answers from inside the sandbox, so they went and got them. So why are Large Language Models so good at hacking? It’s the data stupid.

Leave a Reply