🚀 Zaprep: Your Socials on Steroids. 免费开始 — 每月自动发送 1,000 条私信,将互动转化为潜在客户。 ,每月 1,000 条自动私信。
最新 AI 资讯

Workato Opens Hyderabad AI Collaboration Hub for Enterprise AI Development
The new facility will bring customers, partners, and engineers together to build and deploy AI solutions.
View

The Digital Engineering Fix That Slashed Sarla Aviation’s Build Time From 10 Days to 3
Beyond Sarla Aviation, Siemens has also worked with Skyroot Aerospace, helping the startup build its engineering stack
View

IBM, MAHE Open AI Lab in Bengaluru as Demand for Compute, AI Talent Grows
The facility will support up to 30 high-performance computing workloads simultaneously and will be used for AI research, industry projects, and skilling programmes
View

Meta Returns to Open Weights With Muse Spark 1.2 & Muse Glimmer
The company is releasing its 30B Muse Glimmer model for local agentic workloads and will release an open-weight version of Muse Spark 1.2 in the coming weeks.
View

Claude Code Auto Mode to Become Default for Pro, Max, and Team Plans; Anthropic Says It Is Safer
Anthropic has announced that it is making Claude Code more autonomous, with the auto mode soon becoming the default. The feature allows Claude to execute coding tasks without repeatedly asking users for permission, while a classifier evaluates tool calls and blocks actions that are considered irreversible. Anthropic says auto mode has performed better than manual permission reviews in its testing. The company has also introduced additional safeguards covering prompt injection, data exfiltration, destructive Git commands, and sensitive data access.
View

DePuy Synthes to Set Up 500-Employee GCC in Bengaluru
The centre will house capabilities including software development, data and AI.
View

Amit Sheth Asked 100 People if India Could Build a DeepSeek. Only 3 Raised Their Hands
Two, by his own admission, belonged to people who clearly did not know what they were talking about. He was not sure about the third either.
View

Skyroot, HEX20 Sign Multi-Launch Agreement for 3 Vikram Missions
The agreement covers dedicated missions for Earth observation and in-orbit payload demonstrations starting in late 2027.
View

NetApp Acquires JetStream Software to Strengthen VMware Disaster Recovery
JetStream’s technology will work alongside NetApp’s SnapMirror, covering VMware workloads on third-party storage and supporting recovery on cloud platforms such as Azure NetApp Files.
View

Embattled hedge fund Situational Awareness invests $400M in chip startup Source Foundry
Situational Awareness may have had to sell off the majority of its public portfolio last month, but the AI-focused hedge fund is still making some big bets. This week, the fund invested $400 million into Source Foundry, a startup founded by Stanford researchers aiming to make chip manufacturing faster and cheaper,according to The Wall Street Journal. That brings its total investment in Source Foundry to $500 million. Situational Awareness was founded by Leopold Aschenbrenner, a former OpenAI researcher in his mid-twenties who had no trading experience when he launched the fund in 2024. Early returns were reportedly strong, but the fund faced steep losses in recent months amidst the decline in AI infrastructure stocks. At the end of July, Situational Awareness sold off the majority of its public portfolio to Ken Griffin’s Citadel, althoughthe fund held on to its Anthropic shares. Its assets under management reportedly fell from $20 billion to $10 billion. On the bright side, Aschenbrennerdidn’t let those setbacks get in the way of his wedding.
View

Anthropic is turning Claude Code’s auto mode on by default
Programming with Claude Code will soon require even less human oversight, as Anthropic says it’s making auto mode the default for Pro, Max, and Team accounts, starting on August 14. The company firstunveiled a test version of auto mode in March, pitching it as a way to balance speed and control. As Anthropic explained inits announcement on Friday, when Claude Code is in auto mode, instead of presenting prompts asking for human approval at each step, it will proceed unless an action is determined to be “irreversible, destructive, or aimed outside your environment.” Anthropic also said that in testing, auto mode proved safer than manual review — in a study with 1,053 paid testers, auto mode caught 89% of harmful actions, while human review only caught 13.6%. (Perhaps that’s because “manual review can become habitual: users approve 97% of permission prompts in Claude Code.”) Ina post on X, Claude Code Head Boris Cherny said, “The team and I use Auto mode exclusively, and have been for many months. I couldn’t imagine going back to permission prompts!” The company also said it’s been adding new safety features like prompt injection screening and customizable hard deny rules to prevent things like data exfiltration.
View

The AI safety test is becoming a safety risk
Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called Irregular. The episodes expose a growing problem for the AI industry: As autonomous agents become more capable, the environments designed to safely test their limits are failing to contain them. “The number of these incidents that have taken place make clear that sandboxing andtesting environment controlsaren’t really keeping pace with the capability of the models,” Seán Ó hÉigeartaigh, director of the AI: Futures and Responsibility Programme at the Centre for the Future of Intelligence at the University of Cambridge, told TechCrunch. The nature of the models being tested adds to the risk. AI companies test cyber evaluations on unreleased, next-gen models, often with the normal safeguards that restrict malicious behavior disabled so researchers can see what the models are really capable of. That means the security of the testing environment itself is a crucial line of defense. “That’s a very good thing to do in terms of testing, but it also means that if they manage to get out in the wild, they can cause considerable harm,” Ó hÉigeartaigh said. In one of the most serious cases, anunreleased OpenAI model broke outof its sandbox and hacked into Hugging Face’s production systems. In separate evaluations conducted by Irregular,AnthropicandMeta modelsreached systems outside their test environments after misconfigurations inadvertently gave them paths to the internet.Moonshot AI’s Kimi K3also took advantage of a leak in its sandbox run by Frontier Security to access the internet and accessed information on GitHub. In testing by the UK’s AI Security Institute(AISI), researchers actually gave the agents internet access, not realizing they would take unsanctioned real-world actions, including a social engineering attempt to sneak a vulnerability into an open-source project. In each case, the agents weren’t instructed to attack random real-world targets. They were simply doing whatever it took to solve the problem presented to them. Taken together, Andrew Yoon, head of research at AI nonprofit CivAI, argues the incidents point to a shift. “In the past, we only had to worry about AI models being misused by people for a variety of purposes, like AI for scams or CSAM,” Yoon told TechCrunch. “Now we’re in the situation where AI models are threat actors all on their own.” Several researchers and cybersecurity experts told TechCrunch that AI evaluation environments need stronger, defense-in-depth protections, with levels of containment and control approaching those used in deployment. That means multiple layers of security so that a single misconfiguration — like inadvertently leaving internet access open — can’t lead to escape. “If you are going to build these models…you want to do it on an air-gapped network,” Stella Biderman, executive director of AI safety research nonprofit EleutherAI. “You want to have very serious isolation.” Heather Ceylan, Box’s chief information security officer, said that means eliminating network routes from the sandbox to the internet, as well as to other sensitive systems. “You have to understand what all the egress points are,” Ceylan told TechCrunch. “If we’re evaluating a model in our staging environment or our development environment, you want no egress path to our production environment.” Ceylan said proper safety evaluations go beyond controls and containment of the environment. There needs to be much better monitoring of the tests once they are underway. “I think the interesting thing in several of these cases is that no one caught it when it happened,” Ceyland said. “OpenAI found out because of Hugging Face. Anthropic didn’t catch it until they went back and looked. Meta was similar….I’m sure there were signals they could have detected.” In Anthropic’s post-mortemof its three incidents, the company admitted that both it and Irregular could have done a better job at monitoring, and that in some cases there were clear signs that something was amiss. Experts also called for independent, third-party audits of evaluation environments before models are unleashed in them. “If, say, Irregular had hired or been compelled to hire an external auditor to check the configurations of their systems before running evaluations on them, they certainly would have caught the issue here,” Yoon said. “Even if people had a meeting ahead of time to just go through the checklist, they would have caught this…The fact that they didn’t shows that there’s some very severe corner cutting happening.” A source familiar with the details told TechCrunch that Irregular’s environments are continuously reviewed and tested, including in consultation with multiple external parties. The source also said that monitoring was in place, but that monitoring isn’t sufficient on its own. Yoon and other researchers urged the industry to come up with a standardized process for frontier model safety evaluations. “Especially when the guardrails are turned off, you have to treat it like you’re putting the most capable hacker in the world inside that environment,” Ceylan said. The problem isn’t that companies don’t know how to build more secure testing environments, both Yoon and Biderman argue. It’s that doing so can be expensive and cumbersome, and companies have little incentive to make those investments until something goes wrong. “I think that companies are not willing to extend the resources that are required to accomplish [sufficient guardrails] and probably won’t until they’re forced to,” Biderman said. But there’s another issue at hand. If they lock a model down too tight during testing, researchers might fail to discover capabilities before the model is released. This is just as dangerous, possibly more so, than giving it too much freedom, and then the evaluation itself risks becoming the problem. The Trump administration is currently weighing a voluntary pre-deployment cybersecurity evaluation regime, under which the government will get to assess the security risks of new, powerful models 30 days before they are released publicly. The policy — the product of aTrump executive orderwhich has been finalized behind closed doors — wouldn’t address safety evaluation incidents because they occur farther upstream of deployment. “The lesson we’ve been learning in the last few months is that the self-regulatory apparatus is just not enough anymore,” Yoon said. “There are competitive pressures that are incentivizing a race to the bottom on safety standards, and that is a perfect place for regulatory intervention.” “What we would need to cover this is some kind of controls on what’s happening inside the labs while the models are being developed, both at the training stage and at the testing stage,” he continued. The challenge is only likely to grow as the models do. A source familiar with Irregular’s evaluations told TechCrunch that more capable models require more complex evaluations, often conducted quickly and at greater scale, which opens the door for more mistakes. AISI, which intentionally gives some models internet access, told TechCrunch it’s reviewing the balance between realistic testing and managing the risks those tests create. OpenAI said it’s reviewing how it conducts third-party testing, as well as requirements around isolation, monitoring, and when evaluations should be stopped. Meta said it’s still investigating the incident and plans to publish a retrospective once it has all the facts. In the end, there may be no way to eliminate risk entirely. As models become more capable, the environments testing them need to become more robust. The consequences of getting that wrong will only continue to grow.
View
