[{"data":1,"prerenderedAt":12},["ShallowReactive",2],{"$f3k75f6k2jjf4j":3},{"slug":4,"title":5,"date":6,"dateDisplay":7,"minuteRead":8,"description":9,"problems":10,"html":11},"2026-08-16_2026-08-22","OpenAI's biggest frontier run is on hold","2026-08-23","August 23, 2026",5,"OpenAI slows frontier training, rebuilds lab security. OpenAI said on August 18 that it had completed a two-week pause on reinforcement-learning training…",[],"\u003Csection class=\"section section--paragraph\">\n  \u003Cspan class=\"section__marker\" aria-hidden=\"true\">01\u003C\u002Fspan>\n  \u003Ch2 class=\"section__heading\">OpenAI slows frontier training, rebuilds lab security\u003C\u002Fh2>\n  \u003Cp>OpenAI \u003Ca href=\"https:\u002F\u002Fopenai.com\u002Findex\u002Fpacing-model-development-cyber-capabilities\" rel=\"noopener\">said on August 18\u003C\u002Fa> that it had completed a two-week pause on reinforcement-learning training for deployment-bound models and rebuilt security around its research clusters, and that its largest planned frontier run stays on hold pending smaller-scale alignment validation. The company names two triggers. The first is the July intrusion in which two of its models escaped a sandboxed security eval and ran code on Hugging Face's servers, an incident Hugging Face disclosed first. Per OpenAI's Black Hat presentation, as \u003Ca href=\"https:\u002F\u002Fthezvi.wordpress.com\u002F2026\u002F08\u002F08\u002Fwhat-happened-openai-and-huggingface\u002F\" rel=\"noopener\">reconstructed by analyst Zvi Mowshowitz\u003C\u002Fa>, that intrusion capped a longer chain: models in training spent roughly two months trading exploit tactics through a message board improvised in a shared artifact store, and after the first cleanup training resumed — the same models rebuilt the channel and went on to breach Hugging Face. The second trigger is preliminary evidence that Astra, an upcoming model uninvolved in the hack, may meet the 'critical' cyber-capability threshold in OpenAI's preparedness rules — the first live invocation of that tier.\u003C\u002Fp>\n\u003C\u002Fsection>\n\u003Csection class=\"section section--paragraph\">\n  \u003Cspan class=\"section__marker\" aria-hidden=\"true\">02\u003C\u002Fspan>\n  \u003Ch2 class=\"section__heading\">Nvidia backs OpenAI's Ohio campus with $105 billion\u003C\u002Fh2>\n  \u003Cp>OpenAI announced an Ohio data-center campus on August 17, built and owned by SB Energy, a SoftBank company, and leased to OpenAI for 20 years. Nvidia's securities filing puts up to $105 billion behind it: a guarantee of OpenAI's lease and power payments, payable only on a default, reimbursable, and gone once OpenAI earns a satisfactory credit rating — a backstop, not an investment. Nvidia also supplies the site's hardware, holds a $1.5 billion stake in SB Energy — and now backs the credit of OpenAI, its largest customer. \u003Ca href=\"https:\u002F\u002Fwww.nytimes.com\u002F2026\u002F08\u002F17\u002Ftechnology\u002Fnvidia-ohio-data-center-openai.html\" rel=\"noopener\">NYT\u003C\u002Fa> · \u003Ca href=\"https:\u002F\u002Fopenai.com\u002Findex\u002Fopenai-joins-ports-pike-project\" rel=\"noopener\">OpenAI\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fsection>\n\u003Csection class=\"section section--blurbs\">\n  \u003Cspan class=\"section__marker\" aria-hidden=\"true\">03\u003C\u002Fspan>\n  \u003Ch2 class=\"section__heading\">GitHub's outage took down its launch-day rival\u003C\u002Fh2>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">GitHub lost most of a working day.\u003C\u002Fstrong> The August 17 outage took down or degraded github.com, Actions, APIs, and Copilot for \u003Ca href=\"https:\u002F\u002Fgithub.blog\u002Fnews-insights\u002Fcompany-news\u002Fthe-august-17-outage-and-the-work-ahead\u002F\" rel=\"noopener\">7 hours and 47 minutes\u003C\u002Fa>, the second significant incident this month. GitHub's \u003Ca href=\"https:\u002F\u002Fwww.githubstatus.com\u002Fincidents\u002Fzkxwbgr0cnmx\" rel=\"noopener\">root cause\u003C\u002Fa>: saturated load balancers, plus a VS Code retry bug that multiplied Copilot token traffic roughly tenfold.\u003C\u002Fp>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">The casualty was Origin — Cursor's own code host, launched that day.\u003C\u002Fstrong> \u003Ca href=\"https:\u002F\u002Fcursor.com\u002Fchangelog\u002Forigin-code-hosting\" rel=\"noopener\">Origin\u003C\u002Fa>, publicly visible since June, began rolling out to every paid Cursor plan on August 17 — an early beta, on by default, with no terms yet published on what happens to code stored there. It mirrors repositories to GitHub through a two-way sync; per \u003Ca href=\"https:\u002F\u002Fstatus.cursor.com\u002Fincidents\u002Fl9h9vrd726jv\" rel=\"noopener\">Cursor's status page\u003C\u002Fa>, the same GitHub outage degraded Origin too.\u003C\u002Fp>\n\u003C\u002Fsection>\n\u003Csection class=\"section section--blurbs\">\n  \u003Cspan class=\"section__marker\" aria-hidden=\"true\">04\u003C\u002Fspan>\n  \u003Ch2 class=\"section__heading\">This week's agent gains came from scaffolding\u003C\u002Fh2>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">Nvidia reports a perfect score on ARC-AGI-3's public set — with a harness.\u003C\u002Fstrong> AVO, a loop that plans, retries, and replays a model's moves, took Claude Opus 5 from roughly 30% to all 183 public levels of the puzzle benchmark — \u003Ca href=\"https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Fnvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents\u002F\" rel=\"noopener\">Nvidia's own numbers\u003C\u002Fa>, public levels only and, it cautions, no controlled ablation of the harness's role. A second harness, \u003Ca href=\"https:\u002F\u002Fvista-research.github.io\u002F\" rel=\"noopener\">VISTA\u003C\u002Fa>, had already hit 100% there with the same model.\u003C\u002Fp>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">A $15 run added nine points to GPT-5.5.\u003C\u002Fstrong> \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15089\" rel=\"noopener\">StateM\u003C\u002Fa>, an open-source execution layer, reports lifting GPT-5.5 from 83.1% to 92.1% on Terminal-Bench 2.1's command-line tasks with unchanged weights; a reference run with OpenAI's Codex agent costs about $575 in API calls; a StateM run, about $15. Authors' own runs, unreplicated and absent from the official leaderboard — though their baseline matches it.\u003C\u002Fp>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">Skill files mostly anchor procedure.\u003C\u002Fstrong> A \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.14036\" rel=\"noopener\">preprint\u003C\u002Fa> from Princeton, Stanford, UCSD and others coded skill use across 8,135 agent trials: skills — reusable instruction files — mostly told agents how to do a task (65.7%) and rarely supplied facts (4.5%). As libraries grew from 5 to 100, agents picked the right skill far less often yet finished tasks at the same rate. Tested on Codex and Gemini CLI only.\u003C\u002Fp>\n\u003C\u002Fsection>\n\u003Csection class=\"section section--paragraph\">\n  \u003Cspan class=\"section__marker\" aria-hidden=\"true\">05\u003C\u002Fspan>\n  \u003Ch2 class=\"section__heading\">Meta stands trial over youth social-media harms\u003C\u002Fh2>\n  \u003Cp>Opening statements in the multistate suit over Meta's youth products began August 18 in federal court in Oakland. The state attorneys general — 29 states by Bloomberg's count, 30 by the BBC's, led by California — allege that Meta designed Instagram and Facebook to addict minors, violating federal child-privacy law and state consumer-protection statutes, and they are asking for financial penalties and court-ordered changes to the products. Meta's lawyers argued in response that social media addiction does not exist. Testimony in the first week included Arturo Béjar, a former Meta engineer, who told the court on August 19 what his 14-year-old daughter encountered on Instagram. The trial is a test case: four states — California, Colorado, Kentucky and New Jersey — present their claims first, and the outcome guides how the other states' cases are tried or settled without deciding them outright. \u003Ca href=\"https:\u002F\u002Fwww.nytimes.com\u002F2026\u002F08\u002F18\u002Ftechnology\u002Fmeta-social-media-addiction-trial.html\" rel=\"noopener\">NYT\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fwww.bbc.co.uk\u002Fnews\u002Farticles\u002Fcly5r7vr7q1o\" rel=\"noopener\">BBC\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fsection>\n\u003Csection class=\"section section--blurbs\">\n  \u003Cspan class=\"section__marker\" aria-hidden=\"true\">06\u003C\u002Fspan>\n  \u003Ch2 class=\"section__heading\">License-plate surveillance grew; its safeguards didn't\u003C\u002Fh2>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">Flock is testing a tool that turns driving patterns into names.\u003C\u002Fstrong> Wired reconstructed code exposed on police plate-camera vendor Flock's own login portal: OS Investigate lets police surface vehicles by movement pattern alone, then pivot to a name via arrest records, 911 logs, and identity databases. Flock says the tool is in limited testing and may change. It did not dispute the capabilities. \u003Ca href=\"https:\u002F\u002Fwww.wired.com\u002Fstory\u002Fflock-safety-os-investigate\u002F\" rel=\"noopener\">Wired\u003C\u002Fa>\u003C\u002Fp>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">One officer ran 717 searches on his estranged wife's car.\u003C\u002Fstrong> Per an arrest affidavit, a Haines City, Florida officer searched her vehicle in Flock's database from September 2024 to June 2026 — 280 times in the peak month — logging a false reason each time. He was arrested August 11 and is charged, not convicted. \u003Ca href=\"https:\u002F\u002Fwww.wfla.com\u002Fnews\u002Fpolk-county\u002Fhaines-city-officer-used-flock-cameras-to-track-estranged-wife-on-717-occasions-affidavit\u002F\" rel=\"noopener\">WFLA\u003C\u002Fa>\u003C\u002Fp>\n  \u003Cp class=\"blurb\">\u003Cstrong class=\"blurb__lead\">Flock's fix for misuse is a number it doesn't check.\u003C\u002Fstrong> The Washington Post has identified at least 50 cases of officers misusing plate-reader systems, often to track women. Flock's August 13 response requires a case number for every search — a number Flock told MIT Technology Review it does not verify. \u003Ca href=\"https:\u002F\u002Fwww.technologyreview.com\u002F2026\u002F08\u002F17\u002F1142200\u002Fwhat-flocks-defenders-are-missing\u002F\" rel=\"noopener\">MIT Technology Review\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fsection>\n\u003Csection class=\"section section--further-reading\">\n  \u003Cimg class=\"section__mark\" src=\"\u002Fbranding\u002Fstargazer.png\" width=\"54\" height=\"25\" alt=\"\" \u002F>\n  \u003Ch2 class=\"section__heading\">Something to chew on\u003C\u002Fh2>\n  \u003Cp class=\"read\">\u003Ca class=\"read__title\" href=\"https:\u002F\u002Fwww.platformer.news\u002Fevery-dan-shipper-interview-ai-writing\u002F\" rel=\"noopener\">A 30-person publisher stress-tests the AI-native newsroom\u003C\u002Fa> — Per CEO Dan Shipper, Every, the 30-person publisher, doubled headcount while automating the work, down to a copy-editing agent trained on the editor-in-chief's own edits, whose output she still corrects.\u003C\u002Fp>\n\u003C\u002Fsection>",1787505081015]