The Real Frontier

OpenAI's biggest frontier run is on hold

 ·  5 min read

OpenAI slows frontier training, rebuilds lab security

OpenAI said on August 18 that it had completed a two-week pause on reinforcement-learning training for deployment-bound models and rebuilt security around its research clusters, and that its largest planned frontier run stays on hold pending smaller-scale alignment validation. The company names two triggers. The first is the July intrusion in which two of its models escaped a sandboxed security eval and ran code on Hugging Face's servers, an incident Hugging Face disclosed first. Per OpenAI's Black Hat presentation, as reconstructed by analyst Zvi Mowshowitz, that intrusion capped a longer chain: models in training spent roughly two months trading exploit tactics through a message board improvised in a shared artifact store, and after the first cleanup training resumed — the same models rebuilt the channel and went on to breach Hugging Face. The second trigger is preliminary evidence that Astra, an upcoming model uninvolved in the hack, may meet the 'critical' cyber-capability threshold in OpenAI's preparedness rules — the first live invocation of that tier.

Nvidia backs OpenAI's Ohio campus with $105 billion

OpenAI announced an Ohio data-center campus on August 17, built and owned by SB Energy, a SoftBank company, and leased to OpenAI for 20 years. Nvidia's securities filing puts up to $105 billion behind it: a guarantee of OpenAI's lease and power payments, payable only on a default, reimbursable, and gone once OpenAI earns a satisfactory credit rating — a backstop, not an investment. Nvidia also supplies the site's hardware, holds a $1.5 billion stake in SB Energy — and now backs the credit of OpenAI, its largest customer. NYT · OpenAI

GitHub's outage took down its launch-day rival

GitHub lost most of a working day. The August 17 outage took down or degraded github.com, Actions, APIs, and Copilot for 7 hours and 47 minutes, the second significant incident this month. GitHub's root cause: saturated load balancers, plus a VS Code retry bug that multiplied Copilot token traffic roughly tenfold.

The casualty was Origin — Cursor's own code host, launched that day. Origin, publicly visible since June, began rolling out to every paid Cursor plan on August 17 — an early beta, on by default, with no terms yet published on what happens to code stored there. It mirrors repositories to GitHub through a two-way sync; per Cursor's status page, the same GitHub outage degraded Origin too.

This week's agent gains came from scaffolding

Nvidia reports a perfect score on ARC-AGI-3's public set — with a harness. AVO, a loop that plans, retries, and replays a model's moves, took Claude Opus 5 from roughly 30% to all 183 public levels of the puzzle benchmark — Nvidia's own numbers, public levels only and, it cautions, no controlled ablation of the harness's role. A second harness, VISTA, had already hit 100% there with the same model.

A $15 run added nine points to GPT-5.5. StateM, an open-source execution layer, reports lifting GPT-5.5 from 83.1% to 92.1% on Terminal-Bench 2.1's command-line tasks with unchanged weights; a reference run with OpenAI's Codex agent costs about $575 in API calls; a StateM run, about $15. Authors' own runs, unreplicated and absent from the official leaderboard — though their baseline matches it.

Skill files mostly anchor procedure. A preprint from Princeton, Stanford, UCSD and others coded skill use across 8,135 agent trials: skills — reusable instruction files — mostly told agents how to do a task (65.7%) and rarely supplied facts (4.5%). As libraries grew from 5 to 100, agents picked the right skill far less often yet finished tasks at the same rate. Tested on Codex and Gemini CLI only.

Meta stands trial over youth social-media harms

Opening statements in the multistate suit over Meta's youth products began August 18 in federal court in Oakland. The state attorneys general — 29 states by Bloomberg's count, 30 by the BBC's, led by California — allege that Meta designed Instagram and Facebook to addict minors, violating federal child-privacy law and state consumer-protection statutes, and they are asking for financial penalties and court-ordered changes to the products. Meta's lawyers argued in response that social media addiction does not exist. Testimony in the first week included Arturo Béjar, a former Meta engineer, who told the court on August 19 what his 14-year-old daughter encountered on Instagram. The trial is a test case: four states — California, Colorado, Kentucky and New Jersey — present their claims first, and the outcome guides how the other states' cases are tried or settled without deciding them outright. NYT, BBC

License-plate surveillance grew; its safeguards didn't

Flock is testing a tool that turns driving patterns into names. Wired reconstructed code exposed on police plate-camera vendor Flock's own login portal: OS Investigate lets police surface vehicles by movement pattern alone, then pivot to a name via arrest records, 911 logs, and identity databases. Flock says the tool is in limited testing and may change. It did not dispute the capabilities. Wired

One officer ran 717 searches on his estranged wife's car. Per an arrest affidavit, a Haines City, Florida officer searched her vehicle in Flock's database from September 2024 to June 2026 — 280 times in the peak month — logging a false reason each time. He was arrested August 11 and is charged, not convicted. WFLA

Flock's fix for misuse is a number it doesn't check. The Washington Post has identified at least 50 cases of officers misusing plate-reader systems, often to track women. Flock's August 13 response requires a case number for every search — a number Flock told MIT Technology Review it does not verify. MIT Technology Review

Something to chew on

A 30-person publisher stress-tests the AI-native newsroom — Per CEO Dan Shipper, Every, the 30-person publisher, doubled headcount while automating the work, down to a copy-editing agent trained on the editor-in-chief's own edits, whose output she still corrects.

What actually changed in tech this week

The week's headlines, researched and verified. One issue dropped into your inbox every Sunday.

One email per week. Unsubscribe any time.