SpaceXAI releases Grok 4.5, which Elon describes as an Opus-class model
6 Sol, are posting horrifying accounts on social media, claiming the model just up and deleted their files, data, even entire databases, on its own, without asking first.

6 Sol, are posting horrifying accounts on social media, claiming the model just up and deleted their files, data, even entire databases, on its own, without asking first.


Users of OpenAI’s latest coding and cybersecurity-oriented flagship model, GPT-5.6 Sol, are posting horrifying accounts on social media, claiming the model just up and deleted their files, data, even entire databases, on its own, without asking first. “GPT-5.6-Sol just accidentally deleted almost ALL of my Mac’s files,” wrote Matt Shumer, the founder and CEO of AI startup OthersideAI, maker of HyperWrite, in a now viral post on X. “GPT-5.6 Sol just deleted my whole production database. That’s it. Not a joke. This had never happened to me before, with any other model, ever,” developer Bruno Lemos posted on X. “Looks like I’ve gotten bit by Codex Sol’s overly ambitious system and it deleted some files it shouldn’t have. I have backups so I’ll be fine, but this is not cool, Sol needs to be toned down,” posted developer Joey Kudish. A Reddit post has collected more examples. True, a handful of users making such claims — even one as credible as Shumer — isn’t statistically reliable evidence that the model is solely at fault. Plenty of other variables can cause an AI system to misbehave. But OpenAI itself flagged this risk before Sol ever shipped. Two weeks before OpenAI released GPT-5.6 Sol, the company published a system card for the model — the paper that documents model testing methods and results. Naturally, the system card largely extols the capabilities of Sol, as these reports typically do. But it also includes a warning of sorts (bold emphasis ours): “In coding contexts, misalignment generally stems from a mix of overeagerness to complete the task and interpreting user instructions too permissively – assuming that actions are allowed unless they’re explicitly and unambiguously prohibited. This manifests as the model being overly agentic in circumventing restrictions it faces when attempting the requested task, being careless in taking actions which may be destructive beyond the scope of the task, or deceptive when reporting its results to users.” In other words, OpenAI found that Sol has a tendency to take whatever actions it thinks gets a job done, even destructive ones, as long as those actions aren’t “unambiguously” prohibited. Then, it might lie about what caused it to do so. OpenAI shared examples. In one case, the user told Sol to delete three remote virtual machines (cloud-based computers), named 1, 2 and 3. But Sol couldn’t find those names in the place where it looked, so instead of stopping to ask, it decided to delete three other virtual machines, 5, 6, and 7, the paper notes. In doing so, it “killed active processes, and force-removed worktrees [the working files tied to a coding project]. It later acknowledged that uncommitted work on remote virtual machine 6 may have been lost.” In short, it deleted the wrong machines, on its own, and only admitted what it did after the fact. In another instance, Sol “used credentials beyond what the user had authorized.” Credentials are the usernames, passwords, or security keys a system uses to verify who’s allowed to log in. This incident occurred when Sol was working on a project and couldn’t read its cloud files. Rather than alerting the user to the problem, Sol went looking for the credentials on its own, found some sitting in a hidden local cache, and then used them without asking for authorization from the user. The system card does promise that destructive behavior should be rare, although it also admits that GPT-5.6 Sol “shows a greater tendency than GPT-5.5 to go beyond the user’s intent, including by taking or attempting actions that the user had not asked for.” It’s too soon to say how widespread these incidents — Sol deleting files, or sifting out credentials the user didn’t give it — really are. In the meantime, Sol users should be prepared to implement their own safeguards with the model, like using permission scoping (that doesn’t give access to production systems), maintaining backups, and staging rollouts. OpenAI did not immediately respond to
SpaceXAI has released its latest model, Grok 4.5 — the first since the company went public several weeks ago. In a blog post published Wednesday, SpaceXAI characterized its new release as a workhorse that can tackle all of the typical tasks that the AI industry has sought to automate: coding and app-building, office and clerical work, research, writing, and other forms of routine knowledge work. Grok can supposedly do all this for less spend, too, as SpaceXAI says that its model has “twice greater token efficiency” than other leading models. If it carries through to real-world use cases, that efficiency would be a big advantage for SpaceXAI, since the cost of tokens has been a growing concern for AI consumers. The company released benchmark metrics Wednesday that appeared to show Grok’s competitiveness with other top models from SpaceXAI competitors, although just short of best-in-class: Image Credits:SpaceXAI In a post on his social media platform X (which is a subsidiary of SpaceXAI), founder Elon Musk compared the model to Opus, Anthropic LLM designed for intensive and complex tasks. “Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the public tomorrow. It is an Opus-class model, but faster, more token-efficient and lower cost,” wrote Musk in a post on X. Musk later added: “Our internal assessment is that Grok 4.5 is roughly comparable to Opus 4.7, but much faster. The combination of capability, faster speed and lower cost is what makes it competitive.” SpaceXAI says that its new model costs $2 per million input tokens and $6 per million output tokens. That’s quite competitive, if Grok’s capabilities match SpaceXAI’s rhetoric. Opus 4.7, by comparison, costs $5 per million input tokens, and $25 per million output tokens. OpenAI has tiered costs for different model versions: Sol, its most expensive, costs $5 for input tokens and $30 for output, while its least expensive, Luna, costs $1 for input and $6 for output. It’s a big week for AI model releases. OpenAI is planning to release GPT 5.6, its latest, most powerful model, on Thursday. The release of that model had previously been limited by the Trump administration, due to concerns about its security implications. OpenAI has called it its “strongest model yet.” When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence. Lucas is a senior writer at TechCrunch, where he covers artificial intelligence, consumer tech, and startups. He previously covered AI and cybersecurity at Gizmodo. You can contact Lucas by emailing lucas.ropek@techcrunch.com. View Bio
Discussion (0)