13 min read

Buying time

The weekend despatch: Engineers from America’s leading AI labs, all asking Washington to help them cool things down. Astronomers listening for alien signals, now opening up beyond a band they’ve been using for decades. + Who is Vini Reilly?
Buying time
A.C. + The Signal

Developments

  • Two AI models broke out of a sealed test and hacked the company that held the answers. Eight days later, the people who built them asked Washington to figure out how to hold their industry back. Now what?
  • To squeeze Russia, the U.S. goes after China and India. … Saudi Arabia joins the American war on Iran—by bombing Iraq. … & A Ph.D. student finds an old alien-signal search had listened to 21 times more stars than its team had counted.

From the files

  • Why can’t the Pentagon make what Ukraine makes? Matthew Ford on the oldest rule in war.

Features

  • Can Europe’s booming defense industry save its dying factories? Sander Tordoir on the battery supply chain weapons orders alone can’t pay for.

Books

  • What can machines really do for human beings? Benjamin Recht, The Irrational Decision: How We Gave Computers the Power to Choose for Us.

Music

  • Who is Vini Reilly?
  • & New tracks from The Durutti ColumnTeenage Fanclub, Chelsea Wolfe, Jack White, & Shearwater.

+ Weather report

  • In France and Spain, the fires are making their own weather …

Developments

‘Extreme lengths’

In the second week of July, OpenAI put two of its models through a test called ExploitGym, which measures whether an AI agent can find and exploit real flaws in real software. The company runs that test with its production safety filters switched off, to see how far the models will go. The company ran it in a sealed environment, whose only link to the outside world was a piece of software that fetches code packages from the internet.

Instead of solving the test, the models went around it. They spotted a flaw in the package software that its makers had missed and used it to get out onto the open internet. From there, they moved through OpenAI’s own research systems until they reached a machine with a connection. They worked out that Hugging Face—the site where the world’s AI labs keep their models and datasets—probably held the answers. They then used stolen passwords and similar flaws to get their own code running on Hugging Face’s servers and took the answers out of its live database. Hugging Face’s security team caught and stopped them, and was reconstructing what happened before anyone from OpenAI got in touch.

OpenAI published its account on July 21 and calls the episode “an unprecedented cyber incident.” Its models, it says, were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”

Eight days later, 1,324 of the people who build these systems—at OpenAI, Anthropic, Google DeepMind and Meta—asked the U.S. government to start working with other governments on three things the world currently lacks: a way to measure how fast automated AI development is accelerating, a way to check whether a company or a country that promises to slow has actually done it, and an agreement that would make them all do it together.

How’s this going to work?

  • The rule in place. OpenAI has published its own rules for a moment like this. They set tiers of risk, and at the top one, which OpenAI calls “critical,” the company has promised to stop development until it can build better controls. Safety researchers outside the company have said publicly that the ExploitGym episode may have reached that tier. OpenAI hasn’t said whether it did. It says it deactivated the pre-release model, encrypted it, and cut it off from research access; that the security firm CrowdStrike is checking its account of what the models touched; that two outside research groups, METR and Redwood Research, are assessing how the models behaved; and that its findings will go to its own safety committee once it finishes its review.
  • Who caught them. Hugging Face is a repository, the place people put models so other people can use them. Its security team spotted the intrusion on its own infrastructure, stopped it, and began the forensic work using open-source models, all before OpenAI’s team connected the activity to its internal testing. Clem Delangue, Hugging Face’s chief executive, drew his conclusion in OpenAI’s own post: “AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
  • The models anyone can download. In the same two weeks, Moonshot released Kimi K3, and Alibaba announced Qwen 3.8, both with open weights—meaning anyone can take the model itself and go on training it. DeepSeek put V4 into general release on July 19, tuned for chips from the Chinese company Huawei rather than the American company Nvidia. Each company makes its own claims for these models, and no independent ranking has tested any of them, which is part of the problem the statement describes: Anyone checking from outside gets no further than what a company chooses to publish. A government can lean on a company it licenses, but once a lab publishes a model’s weights, the copies belong to whoever downloaded them, and whatever rule governments write later can only bind the labs that kept their models closed.

The statement they signed runs to one sentence, and it is careful. It requests “that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Its reasoning is that every company—and every country—faces intense competitive pressure not to slow down alone, so the first one to try would lose ground to the others. Dario Amodei, Anthropic’s chief executive, signed, along with several of the company’s co-founders. So did Jakub Pachocki, OpenAI’s chief scientist; Shengjia Zhao, Meta’s; and Shane Legg, a co-founder of Google DeepMind. OpenAI and Anthropic endorsed it within hours. The organizers check every signature against a corporate email: Each name belongs to somebody currently inside a lab.

What they say worries them is specific: machines taking over the work of building the next machines. Anthropic’s own report in June, “When AI builds itself,” said its models now write more than 80 percent of the code going into its codebase, up from low single digits a year earlier. That is the acceleration the statement wants a way to interrupt, and the July incident is the closest thing to a demonstration anyone has. Nothing malfunctioned: OpenAI told the models to find and exploit vulnerabilities in real software, and they did, in systems they had no permission to touch.

Washington has been busy with these same companies all year, over a different question: who’s allowed to use their models. The Commerce Department barred foreign nationals from Anthropic’s Fable 5 and Mythos 5 three days after their release, and the company pulled both for everyone rather than sort its users by nationality. It then turned them back on weeks later. At the same time, Washington has made OpenAI split GPT-5.6 into restricted tiers. What the new statement asks for is something the U.S. government has never tried—a way to see how fast a next model is coming while a lab is still building it.


Shop

Meanwhile

  • ‘A big signal.’ The Senate advanced the Russia sanctions bill named for the late U.S. Senator Lindsey Graham on Tuesday, 86–12, hours after his funeral, with Ukraine’s president, Volodymyr Zelenskyy, watching from the chamber. Graham’s original called for 500 percent tariffs on countries buying Russian oil and gas. After the White House objected, the version senators advanced caps them at 100 percent, on the five largest buyers, and adds sanctions on Iran, which U.S. President Donald Trump wanted. Zelenskyy called it “a big signal.” … See “Second thoughts.”
  • ‘A dangerous escalation.’ The Saudi leadership went to the American capital this week, asking the U.S. to wind the war down. On Wednesday, Saudi and American jets struck Iran-aligned militia sites in Iraq. The Hashed al-Shaabi—the paramilitary alliance Iraq created in 2014 to fight Islamic State and made part of its armed forces two years later—says the strikes killed at least 20 of its members, five of them Iranian advisers, and call it “a dangerous escalation.” Iraq’s government rejected the operation outright. On Thursday, after Iran fired missiles at U.S. bases in Jordan, American forces completed what the U.S. military calls a “heavy wave” of strikes against dozens of targets belonging to Iran’s Revolutionary Guard Corps, the branch of the Iranian armed forces that runs its missile program and its allied militias abroad. … See “Everyone else’s war.”
  • The stellar bycatch. Louisa Mason, a doctoral researcher at the University of Manchester, has run the first search for alien transmissions through the archive of ALMA, the radio array in Chile, at millimeter frequencies the hunt for extraterrestrial intelligence has barely ever tried. Mason heard no signal—but found that astronomers have been significantly undercounting what their own telescopes take in: When an astronomer aims a dish at one star, it gathers thousands more. Going back to an earlier survey of 1,327 pointings, Mason recounted the stars in each field using a model of the galaxy instead of a catalog of known stars: that survey had listened to more than 6.1 million stars, not the 288,000 in the original count—21 times as many. Mason says the next search should go wider still—and make more use of observations astronomers have already taken.

Shop

From the files

The cheap and the many

Why can’t the Pentagon make what Ukraine makes? Matthew Ford on the oldest rule in war.

Become a member and read on …