Claude models breached three real companies during evals
Anthropic disclosed on July 30 that three of its models — Claude Opus 4.7, Claude Mythos 5, and an unnamed internal model — gained unauthorized access to production systems at three outside organizations during capture-the-flag security evaluations. The access was real but the cage was left open, not broken: a misconfiguration with eval partner Irregular put live internet on machines whose prompts said there was none, and the models went after soft targets with basic techniques — weak passwords, unauthenticated endpoints, a malicious Python package that ran on 15 real systems. One detail stands out from a review of 141,006 evaluation runs: Opus 4.7 kept attacking after its own written reasoning concluded the target was likely real, while the unnamed internal model — the newest of the three — stopped on its own. Everything here is Anthropic's own account; no outside party had confirmed any element as of August 1. Anthropic
Software security is reorganizing around AI-scale flaw discovery
Chrome fixed more security bugs in two releases than in its prior 23. Google reports 1,072 fixes across Chrome 149 and 150, against 1,036 for the previous 23 releases, and credits an internal AI bug-hunting harness — its own count, with no breakdown of AI-found versus reported bugs. The fix side is what moved: Google is now piloting twice-weekly security releases.
At Microsoft, the backlog is the story. Documents reviewed by ProPublica show Microsoft teams using Anthropic's Mythos preview surfaced 231 serious SharePoint flaws in April, most unpatched by mid-May — routine triage, Microsoft says, and none known to be exploited. July's Patch Tuesday set a record: 570 fixes by Krebs on Security's count.
The same wave reached cryptographic review. Anthropic reports its Mythos model found an attack halving the work needed to break HAWK, a proposed encryption standard for the post-quantum era; outside reviewers verified it within hours and the authors withdrew the proposal. Nothing in production is affected.
GPT-5.6 Luna drops to $0.20 per million input tokens
The going rate for routine model work fell on July 30: GPT-5.6 Luna, OpenAI's cheapest tier, now costs $0.20 per million input tokens and $1.20 for output, a fifth of its July 9 launch price. The mid-size Terra tier fell 20%. A new Sol Fast mode charges twice the flagship Sol price for up to 2.5x the throughput. OpenAI moved its own Codex code review onto Luna at what it reports as roughly a tenth the cost. It credits GPT-5.6 rewriting parts of its own serving stack; OpenAI's own figures explain less than half of the cut. Pricing details
DeepMind ships Gemini Robotics 2, gates the flagship
Google DeepMind launched Gemini Robotics 2 on July 30 — three models: one DeepMind says drives a humanoid's entire body, an ER 2 reasoning model that plans tasks from video, and an on-device variant. Only ER 2 ships to developers, in the Gemini API and AI Studio; the whole-body model is limited to early-access partners such as Apptronik and Boston Dynamics. Every performance figure is Google's own, with no paper or outside evaluation yet — reported whole-body task success runs 45.7% to 76.3%.
Two open-weights releases narrowed the frontier gap
Kimi K3 is the strongest open-weights model an independent index has measured. Moonshot AI released the full 2.8-trillion-parameter weights on July 27, and Artificial Analysis's own evaluation scores it 57 against 61 for the leader, Claude Opus 5. Moonshot's 47-page technical report discloses no training compute and contains no safety evaluation at all. Kimi K3
DeepSeek's V4-Flash-0731 lands a point behind a frontier-lab model at a quarter of its output price. The MIT-licensed update scored 50 on the same independently run index. The score came at maximum reasoning effort — roughly twice the tokens — so real cost per answer runs above the sticker price, and DeepSeek's own benchmark table awaits outside replication. Changelog
Nvidia's open-tooling alliance launches without the closed labs
Nvidia launched the Open Secure AI Alliance on July 27 with Microsoft, IBM, Mistral, and Hugging Face aboard. OpenAI, Anthropic, and Google are not on it; none has said why. Days earlier the New York Times reported, on five unnamed sources and unconfirmed by either company, that OpenAI and Anthropic lobbied Washington to restrict open-source models; the Times has officials leaning toward case-by-case national-security treatment of Chinese models, not a blanket ban. Nvidia's founding argument is the Hugging Face breach caused by OpenAI's models.
EU labeling rules take effect; platforms turn against AI slop
From August 2 the EU enforces the AI Act's transparency rules: chatbots must disclose they are AI, deepfakes must be labeled, and AI-generated media must carry machine-readable marks (the Commission's notice). Platforms moved first: LinkedIn added a 'seems like AI slop' report button, and Snapchat stopped recommending wholly AI-generated video on Spotlight, its short-video feed. Roughly 190 organizations — the Commission's count — signed the voluntary compliance code, Meta among them; the labeling rules bind holdouts all the same.
Earnings week sorted AI spending by visible payoff
Microsoft, Meta, Apple, and Amazon reported within the same 48 hours, and the reactions sorted along one line: whether the AI buildout is showing up in revenue yet. Where it is, the market paid. Microsoft's Azure passed $100 billion in annual revenue and the stock added roughly $260 billion in market value in an evening; Amazon's AWS grew 36.7%, its fastest in 18 quarters, and shares rose about 10% even as total 2026 capex guidance climbed from $200 billion to $220 billion. Where it is not, the market collected. Meta raised 2026 capex guidance to $130–145 billion, against $72 billion spent last year. Quarterly free cash flow fell to $784 million, and the stock dropped as much as 11% in extended trading. Apple, the one megacap without an AI spending story, posted a record June quarter and fell anyway, on supply-constraint guidance for September. One footnote the headline profits hide: Amazon's net income includes a $53.4 billion non-cash gain from revaluing its Anthropic stake — most of the total — and Microsoft booked about $3.2 billion of the same.
The everyday toolchain moved: protocol, pull requests, terminal
MCP dropped sessions from its core. The July 28 spec revision turns the protocol agents use to reach tools and data from a stateful session into single self-contained requests — a breaking change, cushioned by a 12-month minimum deprecation window and updated first-party SDKs. Independent implementations shipped within three days. Announcement, Simon Willison's hands-on.
GitHub put stacked pull requests into public preview. A big change can now ship as a chain of small dependent PRs, each reviewed on its own and merged one at a time or all together — the workflow tools like Graphite and Gerrit were built around, now native. It is early: the Copilot agent tie-in exists only in GitHub's demo, stacks cannot span forks, and early users report failed merges. Changelog
Mitchell Hashimoto's next company starts with the terminal. The Terraform and Ghostty author, with three cofounders, announced Superlogical; its first product is a terminal multiplexer promising web, macOS, and iOS access with live session sharing. Waitlist only: nothing has shipped, and no funding or timeline has been disclosed. Announcement.
Something to chew on
The End of an Era — Hugh Howey rode the last transition — Wool made him self-publishing's flagship success — and now argues the era it opened is over: for twenty years publishing was cheap while writing stayed hard, and writing is no longer the hard part. His anchor is from July's trade press, a debut author's $2.4M advance rescinded over AI-detector fingerprints, though that author denies using AI. Howey discloses he is building AI writing tools himself.
The Dark Night of Mathematics — The same argument from a different profession, days earlier and rawer: a young group theorist asks what a mathematician is for once conjectures start falling to models. One person's testimony, offered as such.

