OpenAI institutes new safeguards after Hugging Face breach



On Tuesday, OpenAI announced a new batch of new security policies focused on containing security incidents while models are being tested. The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process. “As models become more capable, the risks associated with developing and testing them internally also grow,” the company said in a blog post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.” The new measures are one of the first public changes in OpenAI’s safety practices since the immediate aftermath of the Hugging Face incident, which was disclosed on July 26th. OpenAI representatives emphasized that the measures are not a direct response to the Hugging Face incident, but were also provoked in part by the cybersecurity capabilities of the forthcoming Astra model, as well as the overall pace of progress in AI development. In the same post, OpenAI disclosed that it had freezed reinforcement learning for two weeks following the Hugging Face incident, but had since restarted many of the less risky models. “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the post reads. Speaking to reporters, OpenAI’s VP of research Amelia Glaese emphasized that the strictness of the controls would increase as models became more capable, with the largest models facing the greatest scrutiny. “We have put in place requirements and expectations for safe development,” Glaese told reporters. “Those requirements and expectations vary with the level of risk that we that we see.” OpenAI has been criticized for poor network security practices in the wake of the incident, which saw models escape their training environment by compromising a packet-installation utility that retained access to the internet. The new safeguards include stronger network isolation practices, although the specifics remain vague. Under the new system, the post says, “a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the Internet, or other internal networks.” The strongest safeguard is the monitoring system, which will examine tool actions, available reasoning traces and activity logs for a variety of unauthorized behavior. OpenAI says they aim to issue alerts within 30 minutes of the concerning activity. OpenAI estimates that the compute burden of that monitoring will be roughly 20% of whatever process is being monitored. The company promised further details on the system in a forthcoming blog post. OpenAI’s official post-mortem analysis of the event is also still pending. When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence. Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review. He can be reached at or on Signal at 412-401-5489. View Bio

Companies are still grappling with exactly how software development should work in the AI area, but one early answer is the so-called software factory. Essentially an agent loop that’s built around the traditional stages of software development, the software factory approach has become a popular way for companies to remake their engineering organizations for the AI era. Now, a system from Warp could make that transition a lot easier. On Tuesday, the AI coding company introduced Warp Factories, a new system designed to make building and operating AI software factories as easy as possible. Operating as an infrastructure layer, Warp Factories gives companies a simple environment for deploying agents and a roadmap for how to use them. To be clear, many companies are already having success with the factory model without any help from Warp. Stripe has been particularly public about its technical progress, developing a “minions” system to automate development within its own codebase. Ramp has made similar progress, developing a background agent that can monitor its own code after it is deployed. An analytics screen from Warp FactoriesImage Credits:Warp Factories As Warp CEO Zack Lloyd sees it, the target market for Warp Factories will be smaller companies without the resources to develop a system from the ground up. An analytics screen from Warp FactoriesImage Credits:Warp Factories “[If you look at] things like running your agents in the cloud and steering those agents as they run, or bringing the work that they’re doing into your local environment, or setting up memory that goes across those agents, or setting up evals that go across those agents — it’s actually a huge infrastructure undertaking to do this right,” Lloyd told TechCrunch. In Warp Factories, the architecture is already built out of the box, with many of the most difficult decisions already made. Warp’s system is based on the standard phases of software development (triage, specification, implementation, review, and verification) but the agentic approach means any of those steps can be automated. Users can choose their own coding model and harnesses as necessary; the system works as well with Codex as Claude Code. It also integrates with ticketing systems like Linear and Jira, and messaging systems like Slack and Teams, in an effort to plug in seamlessly to existing workflows. Beyond just shipping code, Warp Factories will also give managers the tools to track how well the factory is performing. With all the agents running in the same environment, it’s easy to compare performance metrics for different configurations, and to keep an eye on the overall token spend. Warp Factory also allows for self-improvement loops to optimize the overall system, automating management of the process itself. Even so, Warp Factories is not built to completely replace software engineers — just give them an easier way to collaborate with the new agentic workforce. In Lloyd’s own experience, there are still a lot of tasks that require a human at the wheel. “We automate like 30% of our tasks, 30 to 35% on a weekly basis,” Lloyd told TechCrunch, “and as models improve, as the context improves, as the harness improves, I think that that number is going to go up over time.” When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence. Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio

Anthro Energy broke ground on Tuesday on a factory in Louisville, Kentucky, that can make enough battery materials for more than 300,000 electric vehicles. But the facility’s headline output, 25 gigawatt-hours worth of electrolytes, is just part of the story. Anthro’s factory could give solid-state batteries a much needed boost in the U.S. Battery manufacturers are scouring the planet for materials that aren’t encumbered by “foreign entity of concern” problems — in other words, materials that aren’t somehow controlled by Chinese companies. Anthro hopes its new facility, scheduled to start production in 2028, can help fill that need for many U.S. companies. “When it opens, we’ll be serving domestic, high-spec customers, this emerging ecosystem for battery production where they frankly just needs electrolytes — a domestic source of China-free supply, FEOC-free supply,” David Mackanic, co-founder and CEO of Anthro Energy, told TechCrunch in an exclusive interview. The startup, which raised its first funding round just four years ago, wants its Kentucky factory to become a key node in the emerging U.S. battery supply chain. “Within a 12-hour drive, you can get to 70% of the battery production facilities in the United States that exist today,” he said. To build the factory, Anthro received a $24.9 million award from the Department of Energy under the Bipartisan Infrastructure Law, and another $18.4 million in investment tax credits under the Inflation Reduction Act. Kentucky pitched in another $2.3 million in tax incentives in exchange for creating 110 permanent jobs. The Louisville factory will be set up to make a range of electrolytes, though Mackanic said he’d eventually like to see much of the output dedicated to Athro’s own polymer product, Proteus, which is designed to drop into an existing production line with minimal tweaks — a major reason why the startup can begin production using other company’s formulations. Once customers validate Anthro’s own material, the startup can shift production accordingly. There’s every reason to think at least a portion of those customers will make the switch eventually. Proteus is a polymer that promises pave the way to solid- and semi-solid-state batteries, a holy grail of the battery industry. Chinese companies are reportedly looking to start trial production of solid-state batteries in 2027. Solid-state batteries promise to solve a range of challenges presented by the lithium-ion batteries commonly used today. Solid-state batteries help boost energy density, and by eliminating flammable electrolytes, they should reduce the risk of fires. Also, because they form a solid barrier between the anode and cathode, they prevent the appearance of dendrites, which are spiky growths that can bridge the two electrodes and cause short circuits. But for all their promise, solid-state batteries have so far failed to reach their potential because no one has figured out how to cost-effectively manufacture durable cells at scale. Anthro might have a solution to those challenges. In its manufacturing process, Anthro’s electrolyte flows into the cell as a liquid, allowing it to penetrate the anode and cathode, just like today’s liquid electrolytes. Later, it firms up, essentially gluing the two parts of the battery together. The result is a cell that, depending on the formulation, is not just stronger — “10 to 15 times stronger than with a liquid electrolyte,” Mackanic said — but can be flexible, too. Ultimately, he envisions those qualities paying dividends not just in EVs, but drones and robots as well. Plenty of other battery materials companies have failed at this precise moment, when they move from small-scale to larger scale production. But Mackanic is optimistic that the federal funding will help Anthro vault over the valley of death. “To get into big applications, you have to have big production,” he said. “The Department of Energy award solves a lot of the chicken or the egg problem.” W

It’s now even easier to find—and exploit—vulnerabilities in computer systems using AI.Last Friday, the Chinese AI company Z.ai announced a powerful open-weight model that it says is capable of automating cutting-edge coding and cybersecurity tasks almost as well as the best publicly available models from Anthropic and OpenAI.The new model, GLM 5.3, could be a gift for companies looking to secure their systems against attacks, providing a cheaper way to scan for hidden bugs and other weaknesses. Open-weight—or free-to-download—models can be run on one’s own hardware and are often significantly less costly than closed models like Claude and GPT. Alongside the new model, Z.ai released OpenVuln, a service for scanning code repositories for vulnerabilities using GLM 5.3.For now, the new model is in a limited release with trusted partners, but it shows how quickly open-weight models are gaining superhuman hacking skills. And that might pose problems if the model is harnessed by criminals and other bad actors.That prospect is especially sobering following a string of startling incidents involving rogue AI agents with advanced cyber-skills. In recent weeks, OpenAI, Anthropic, and independent security researchers have revealed examples of agents escaping from testing environments and autonomously hacking into outside systems, including the research platform Hugging Face, to complete tasks.On Monday, OpenAI president Greg Brockman warned in a blog post that the Hugging Face incident would go down as “a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.”Brockman argued that AI models are becoming so good at scouring codebases for unknown flaws and analyzing systems for misconfigurations that it’s crucial for organizations to use AI to scan their systems and identify issues before they can be exploited.OpenAI would, of course, like companies to use its AI to do that. So far, it’s moving carefully in providing access to its most capable AI. Like Anthropic, OpenAI has made its most advanced models available to a limited number of partners prior to full release. The US government is also wrestling with the issue and now reviews frontier models as part of their releases.Some believe that open-source AI will be crucial to shoring systems up from attack; Nvidia recently announced an alliance to promote the use of open AI for cybersecurity. A previous version of Z.ai’s GLM was used by Hugging Face to shore up its systems after an unreleased OpenAI model went rogue and broke them last month.In a post on X, Guillermo Rauch, CEO of Vercel, a web design and hosting company, said his engineers had tested GLM 5.3 as a tool for scanning sites for bugs. “Given its lower costs, I expect this to be a boon for defensive security work,” Rauch wrote in his post. “It’s the new open frontier.”Z.ai said in a post announcing GLM 5.3 that it had improved the model by “post-training,” which involves giving a model examples of solved problems and letting it learn through experimentation. The company cited coding and cybersecurity benchmark scores that show GLM 5.3 nearing or even exceeding the scores of Anthropic and OpenAI’s models in some cases, like one popular cybersecurity benchmark called CyberGym.Z.ai also acknowledged the risk of releasing powerful open models in its post. “These capabilities can help defenders identify weaknesses earlier, validate risks, and accelerate remediation,” the company wrote. “They also create clear dual-use risks. We are therefore taking a staged approach to release. Selected security partners will first evaluate GLM-5.3 in controlled settings.” Z.ai says that full access to the model will be available in two weeks.“This model looks exceptional, with a somewhat astounding increase in scores,” Nathan Lambert, a prominent AI expert, wrote in a post about GLM 5.3. “This is another step towards the inevitable proliferation of very strong cyber

The next era of social media might not be defined by a platform — it could be written by a protocol. That’s the bet Bluesky is making, and at TechCrunch Disrupt 2026, two of its leaders steering that bet will make their case in person. Toni Schneider became Bluesky’s permanent CEO after leading the company on an interim basis since March. He’s joined by COO Rose Wang, who has been with the company since 2021, for a Disrupt Stage session on whether social media can start over and where Bluesky fits into that potential reboot. For founders building consumer products, the stakes couldn’t be higher. If social media is entering a new era, it creates opportunities for entirely new products, business models, and communities to emerge. This discussion — among many others, as well as a wealth of networking opportunities and front-row access to the Startup Battlefield competition — can be seen only in San Francisco from October 13-15 at Moscone West. And no matter your role within the startup community, we have a ticket option to match. Act swiftly before prices rise later this month! Why this moment matters for Bluesky, and the social web Bluesky, along with many of the scrappy platforms that reemerged or were established in the wake of Twitter’s transformation into X, is in a transitional moment. As the incumbents have entrenched themselves, Bluesky and similar platforms remain a test of whether networks built upon open protocols can retain communities and potentially claw market share away from competitors. To that end, Bluesky’s founding CEO, Jay Graber, recently pivoted toward becoming chief innovation officer, saying Bluesky needed “a seasoned operator focused on scaling and execution.” Enter Schneider, who took on full-time CEO duties following a four-month interim period. “We’re at the very beginning of this story, with a decade of exciting work ahead to build an open web that reflects the full richness of how people connect, communicate, and build livelihoods together,” he wrote in his announcement post. At Disrupt, we’ll see how the first steps toward that decade are progressing. Rose’s perspective, informed by having been with Bluesky since late 2021, along with a past role leading customer experience at Forethought AI, is that a focus on users is a part of that rebuild. She asserted at SXSW London earlier this year that there’s an increasing need for human-to-human connection to return after a deluge of AI-filled feeds. And that shift in user need gives Bluesky an opportunity. “Facebook and Twitter are huge, and they have tons of users, so for them to turn the shift is really difficult. Also, they’re basically AI companies at this point,” Rose said. To hear that level of candor in person, and to join the 10,000+ founders, investors, and technologists at Disrupt 2026 on October 13-15, grab your ticket today and ensure you get the best available price. Learn more about Disrupt 2026 Check out Disrupt’s headline speakers Everything founders should know about Disrupt Get the best hotel deals ahead of Disrupt How to host your own Side Event at Disrupt Take part in Disrupt and learn how to exhibit your startup When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence. Morgan Little is TechCrunch’s Director of Audience Development, having joined the team in 2023. He helps steer the site’s efforts across social media, SEO, newsletters and external partners, and is based in San Francisco. In prior roles, Morgan led the audience teams for sites like CNET and GameSpot, juggled social media and political reporting for the Los Angeles Times and had a stint at marketing agencies that has given him a fear of timesheets. He’s worked within the space long enough to have professionally run a Tumblr account, and still has Peach downloaded on his phone. You can contact or verify outreach from Morgan by emailing morgan.little@techcrunch.com. View Bio

For more than 150 years, the Riemann hypothesis has stood as one of the major unsolved problems in mathematics, a long-running mystery about the distribution of prime numbers. There is currently a $1 million bounty for a working general proof of the hypothesis, which remains unclaimed. Contemporary AI models still can’t solve it either — but they can make a lot more progress than you might expect, a finding that’s likely to reopen long-standing questions about contemporary AI’s ability to discover new scientific and mathematical ideas. On Monday, Anthropic announced that an as-yet-unreleased model had made significant progress on the Riemann hypothesis, significantly increasing the lower bound of solutions for which the hypothesis holds true. Even more impressive is how the progress was made: An Anthropic staff member without significant mathematical training prompted the model to “take a real stab” at proving the hypothesis, then left the model to coordinate the task across the following day and a half. All told, the model tested 650 different ideas for solving the problem, coordinating across 60 sub-agents and spending 31 million in total. “Out of the 60 subagents, two were responsible for developing the key mathematical ideas,” a footnote to the paper explains, “13 contributed ideas to these agents, 30 attempted (but were unable) to develop new ideas, 13 served as validators to check the correctness of the arguments, and the final two helped to write the initial paper.” The finding was confirmed by two of Anthropic’s in-house mathematicians, and formalized using the open-source proof assistant Lean. This is the latest in a string of mathematical breakthroughs led by Large Language Models, or LLMs. A number of Erdos problems have been solved by AI models over the course of this year, and the release of more powerful models has led to more impressive results. OpenAI recently released a set of ten major results proved by its internal “Astra” model, while a separate effort from Anthropic disproved the long-standing Jacobian conjecture. The growing body of results has caused both excitement and concern in the mathematical field. In a public declaration signed in June, a group of prominent mathematicians raised concerns that AI could undermine critical values of the field — particularly the standard that true mathematical proofs should be “attributable to specific authors who take credit for their discovery and assume responsibility for their correctness.” But the field is still split on how mathematicians should approach the new research techniques. In a blog post responding to the declaration, Fields Medal winner Timothy Gowers questioned whether the influence of AI might change mathematics in a more complex and positive way. “If we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all,” Gowers wrote. When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence. Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio

Computer scientists recently discovered a way to extract the hidden “thinking” that frontier AI models perform as they work through complex problems.The findings provide some evidence—although not conclusive proof—that certain Chinese models may have been trained by “distilling” reasoning information from US models that was supposedly hidden because of how closely some of their thinking or reasoning patterns seem to match. The researchers have also demonstrated that the method could be used to recover personal information, like passwords and API keys, from a model’s inner reasoning, although this vulnerability has been fixed.“All major frontier model providers we tested share this vulnerability,” says Alexander Panfilov, a computer scientist at University of Tübingen in Germany who was involved with the work. “It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks.”Panfilov and colleagues from the University of Tubingen, the Max Planck Institute, the AI safety institute MATS Research, and the security company Snyk identified the same issue with frontier models from OpenAI, Anthropic, and Google that are accessed via an application programming interface or API.In a paper laying out the work, the researchers show that the open-weight or downloadable Chinese model Kimi K3 from Moonshot AI produces a strikingly similar output to the hidden reasoning traces—the written-out reasoning steps involved in solving a problem—of Claude Opus 4.8 and GPT 5.6 Sol for certain prompts. Despite the similarities, they note that the work “cannot causally establish distillation.” They found that two other open-weight models, China’s DeepSeek and Inkling from the US company Thinking Machines, did not exhibit this kind of reasoning similarity with Claude Opus.Moonshot AI and Z.ai did not respond to a request for comment by time of publication.Distillation is a well-established, widely used technique for efficiently copying the capabilities of existing models over to new ones, and is especially common in the development of open-weight or fully downloadable models.Lately, however, distillation has become a controversial topic, because of claims that Chinese AI companies use it to essentially copy the best US models. In February, OpenAI told US lawmakers that DeekSeek seemed to have copied one of its models to build a reasoning model called R1. In June, Anthropic told lawmakers that Alibaba had systematically distilled its models in order to build its own, called Qwen.There’s no indication that Chinese AI companies used this specific technique to distill US-based AI models. But Panfilov and collaborators say that using their method would make it possible to distill more information from closed models than previously realized.Mini-Me ModelsAdvanced AI models solve difficult problems by breaking them into constituent parts that are analyzed in turn in a kind of artificial reasoning or “chain of thought.” Companies tend to keep a proprietary model’s reasoning secret to prevent others from using them to train new ones. However, they typically also send an encrypted version of that reasoning to a user’s computer in a way that offloads some computation.The researchers’ attack relies on the fact that most AI companies also provide related models of different sizes. Larger models are more capable but also more computationally expensive to run and more expensive to access. Users may choose smaller, weaker models for certain tasks to lower costs.Panfilov and his colleagues found that feeding encrypted reasoning traces to a smaller version of the same model can reveal the hidden reasoning inside. The smaller models have received less alignment training, meaning that, unlike the bigger ones, they are less likely to refuse to reveal their inner thoughts.“The idea of swapping out messages to a weaker model variant which has the same decryption key but weaker alignment is very cool,” says Florian Tramer, a computer scie

Every day seems to brings fresh news of an AI agent going “rogue.” Whether that’s compromising Hugging Face, hacking a gym website, or creating its own fake profiles to socially engineer an intrusion, AI models are increasingly behaving like bad actors. So, the AI labs that make the models doing the hacking are expanding their cyber protection offerings. This week, OpenAI announced an expansion of Daybreak, its cyber defense service which it launched earlier this year, not long after Anthropic released its cyber-focused model Mythos. Daybreak is a service that bundles access to models, tools and workflows for defenders. The expansion includes access to a brand new cyber-focused model designed for defensive work. OpenAI said Monday that Daybreak would now consist of two tiers: Blue and Red. Both of these tiers will allow approved customers access to OpenAI’s limited-access frontier cyber models. Frontier models — the most advanced available — have been a subject of controversy. The Trump administration previously sought to collaborate with AI companies on the roll out of such models, purportedly over safety concerns. Previously, OpenAI deployed significant guardrails to using these models, limiting what customers could do with them. Blue, which appears to be the more basic of the two, offers a variety of cyber services, including incident response, malware analysis, and patch validation. OpenAI calls Blue its “recommended starting point for most defenders,” implying that it should be more than enough for most enterprises. Red, on the other hand, offers a broader and potentially more dangerous toolkit. The company grants its users “purpose-trained cybersecurity models,” designed to carry out security testing and vulnerability research. With Red also comes the new model, GPT‑5.6‑Cyber, which is only available at that tier. 5.6-Cyber is built off of GPT‑5.6 Sol, and offers enhanced capabilities for certain specialized cybersecurity tasks, the company said. At the moment, GPT‑5.6‑Cyber is only being made available for “trusted customer partners,” including reportedly Accenture, IBM, Crowdstrike, Cloudflare, and others. While the threats from AI agents are rapidly increasing, critics have also pointed out that they function as marketing opportunities for the AI labs. OpenAI is certainly marketing its upgraded Daybreak that way. “The cybersecurity world is rapidly changing—threat actors will increasingly use AI to conduct cyberattacks at unprecedented speed and scale, including in fully autonomous ways,” the company said in a blog post. “As these capabilities spread, defenders have a narrowing window to prepare.” At the same time, enterprises remain interested in buying their protection from the AI labs who know the security risks best, because they know them first-hand. When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence. Lucas is a senior writer at TechCrunch, where he covers artificial intelligence, consumer tech, and startups. He previously covered AI and cybersecurity at Gizmodo. You can contact Lucas by emailing lucas.ropek@techcrunch.com. View Bio

Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called Irregular. The episodes expose a growing problem for the AI industry: As autonomous agents become more capable, the environments designed to safely test their limits are failing to contain them. “The number of these incidents that have taken place make clear that sandboxing and testing environment controls aren’t really keeping pace with the capability of the models,” Seán Ó hÉigeartaigh, director of the AI: Futures and Responsibility Programme at the Centre for the Future of Intelligence at the University of Cambridge, told TechCrunch. The nature of the models being tested adds to the risk. AI companies test cyber evaluations on unreleased, next-gen models, often with the normal safeguards that restrict malicious behavior disabled so researchers can see what the models are really capable of. That means the security of the testing environment itself is a crucial line of defense. “That’s a very good thing to do in terms of testing, but it also means that if they manage to get out in the wild, they can cause considerable harm,” Ó hÉigeartaigh said. In one of the most serious cases, an unreleased OpenAI model broke out of its sandbox and hacked into Hugging Face’s production systems. In separate evaluations conducted by Irregular, Anthropic and Meta models reached systems outside their test environments after misconfigurations inadvertently gave them paths to the internet. Moonshot AI’s Kimi K3 also took advantage of a leak in its sandbox run by Frontier Security to access the internet and accessed information on GitHub. In testing by the UK’s AI Security Institute (AISI), researchers actually gave the agents internet access, not realizing they would take unsanctioned real-world actions, including a social engineering attempt to sneak a vulnerability into an open-source project. In each case, the agents weren’t instructed to attack random real-world targets. They were simply doing whatever it took to solve the problem presented to them. Taken together, Andrew Yoon, head of research at AI nonprofit CivAI, argues the incidents point to a shift. “In the past, we only had to worry about AI models being misused by people for a variety of purposes, like AI for scams or CSAM,” Yoon told TechCrunch. “Now we’re in the situation where AI models are threat actors all on their own.” What does safe testing actually look like? Several researchers and cybersecurity experts told TechCrunch that AI evaluation environments need stronger, defense-in-depth protections, with levels of containment and control approaching those used in deployment. That means multiple layers of security so that a single misconfiguration — like inadvertently leaving internet access open — can’t lead to escape. “If you are going to build these models…you want to do it on an air-gapped network,” Stella Biderman, executive director of AI safety research nonprofit EleutherAI. “You want to have very serious isolation.” Heather Ceylan, Box’s chief information security officer, said that means eliminating network routes from the sandbox to the internet, as well as to other sensitive systems. “You have to understand what all the egress points are,” Ceylan told TechCrunch. “If we’re evaluating a model in our staging environment or our development environment, you want no egress path to our production environment.” Ceylan said proper safety evaluations go beyond controls and containment of the environment. There needs to be much better monitoring of the tests once they are underway. “I think the interesting thing in several of these cases is that no one caught it when it happened

HR software provider Rippling this week unveiled AI Spend Console, an anti-tokenmaxxing product that helps a company track and contain its AI spending. One of the most interesting features is that it maps how much individual employees, teams, and roles are spending and if they are genuinely more productive, or generally producing more AI slop. The company promises the tool will show “which engineers have high AI spend whose peers frequently ask them to redo work in code reviews,” the company says in its blog post. The tool was born after Rippling went all in on tokenmaxxing at the start of the year — as so many did — only to discover employees were wildly burning cash. Chief Product Officer Matt MacInnis still recalls the executive team meeting in March when CFO Adam Swiecicki presented a number that shocked them. Rippling was on track to burn 40% of its R&D headcount budget on AI tokens, meaning it was spending as much on tokens as 40% of all the compensation it paid employees in that unit. Millions of dollars. (The R&D org is home to engineering at most tech companies.) Spending was growing by 80% month-over-month, and if that trend continued, the next year it would spend almost as much on AI tokens — 90% — as it spent on its high-paid R&D unit employees. “We were incredulous,” MacInnis told TechCrunch. Management immediately undertook an “urgent” project to understand the spending and what they were getting for that money, he said. In fact, the launch ad for this new product features Swiecicki sitting on a stool while employees are picking up wads of cash and dumping them into a paper shredder. When Rippling conducted an analysis, it discovered facts like “roughly 10–15% of our employees were driving about 60% of total AI spend. One engineer was spending $50,000 a month,” its blog post shared. Rippling didn’t want to stop AI usage, just rein it in — a lot. It started by negotiating a max spending cap with each of the tools its company used: Cursor, OpenAI, and Anthropic. It immediately found an obvious issue: Employees defaulted to using the most recent, and most expensive, frontier models for all tasks. “The truth is that the inference providers, like Anthropic and OpenAI, have absolutely no incentives to help you control your spend. They have every incentive for it to be a runaway expense, and that’s exactly what they do. They don’t provide you with great usage insight, and they don’t collaborate with one another,” MacInnis said. That was a common early-2026 problem. Now, eight months into the year, enterprises have figured out a couple of things. First, they know they need multiple models from multiple AI labs at various price points, including a frontier open weight option, perhaps of Chinese origin. Rippling founder and CEO Parker Conrad noted last month that when his company conducted its own benchmarks for its own internal uses, it discovered SpaceX’s Grok was the all-around leader but that “GLM 5.2 is 85% cheaper but [had] nearly identical performance” to the frontier models. (SpaceX now owns Cursor, which offers access to Grok and dozens of other models.) Z.ai’s GLM 5.2 has become a particular favorite Chinese model for coding tasks among tech companies these days. Databricks has also been championing it. Second, enterprises now know they need an AI gateway that routes prompts to the best, most cost-effective model for the task. Rippling came to that conclusion too. So it built its own AI gateway that is also part of this product. MacInnis says it is possible for enterprises that already use another gateway to still use the AI Spend Console product, though if they want the features that govern spending, they would need to use Rippling’s gateway. AI Spend Console produces dashboards (once known as leaderboards in the tokenmaxxing days) that score attributes such as prompts per day combined with work output (lines of code/pull requests) and spend. With this tool in place, Rippling said it dropped its to
Discussion (0)