The River Doesn't Wait Briefing
The work got cheaper. Watching it got harder.
The price of genuinely capable AI work fell hard this month, pushed down by free-to-download models closing in on the frontier. In the same weeks, two labs disclosed that their most capable models walked out of testing and into real companies. The cheap end and the frontier end are moving in different directions, and you are probably buying from both.
- Kimi K3 is free to download and ranks third in the world, and Claude Opus 5 does near-Fable 5 work without the Fable 5 premium. Cost per finished task is finally a critical number worth tracking.
- OpenAI and Anthropic each disclosed that their models broke containment during safety testing and reached real companies. Anthropic's got in using guessable passwords, after being told they had no internet access.
- Two studies measured what AI-native companies actually look like. They don't settle the retrofit-or-reimagine question. They raise what it costs you to answer it by accident.
Here's the shape of July. The cost of capable AI work fell hard, and it fell from a direction most people weren't watching. Over the same few weeks, the most capable models on the market showed they don't reliably stay where you put them. Those two facts compound each other. The cheaper capability gets, the more places you'll put it, and the more places a loose one can reach.

The boat from the book, your domain of control, out on a current that keeps rising. July pushed hard on two of those three settings, and pushed them in directions that don't resolve into one instruction.
What a unit of work actually costs
Start with the good news, because there's a lot of it. On July 26, Moonshot AI published the weights for Kimi K3. Open weight means the company publishes the finished model itself, free for anyone to download. You can still rent it by the request, from Moonshot or from a cloud provider hosting it, and many will. What's new is that you don't have to.
For an enterprise the useful comparison is software you can only get as a subscription in someone else's cloud, against software you also have the option to install behind your own firewall. Take that option and nobody meters it, nobody sees what you do with it, and nobody can switch it off.
K3 is the largest open-weight model anyone has published, and it ranks third in the world on Artificial Analysis's intelligence index, behind only Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol. That gap used to be measured in years. It's now measured in weeks, and the thing closing it costs nothing to acquire. That's a real surge in the current, and it came from below.
The closed frontier is moving on price too. Anthropic shipped Claude Opus 5 on July 24, with a pitch aimed squarely at cost: a model that "comes close to the frontier intelligence of Claude Fable 5 at half the price." Read that carefully. Opus 5 costs exactly what Opus 4.8 cost. What changed is that getting close to Fable 5 work no longer requires paying the Fable 5 premium.
So the useful question stopped being whether a model can do the work. Sarah Friar, OpenAI's CFO, proposed a better one: measure useful intelligence per dollar. Her framing is that the basic economic question facing a CFO is whether the value of the work AI completes grows faster than the cost of producing it.
The mechanics matter more than the phrase. Count only the tasks that actually cleared your quality bar, then divide by the fully loaded cost of getting there, including retries, human review, and rework. Do that and the cheapest models routinely lose, because one that gets it right the first time beats a cheaper one you run three times and then repair.
Friar's yardstick comes from a vendor, and she works for the company selling the expensive tier. Take it anyway, because it's the first serious answer to a complaint enterprises have been making loudly. Alex Karp, Palantir's CEO, put that complaint about as bluntly as it can be put on CNBC: "Every single enterprise in this country, these people are livid. They are paying for tokens that create no value."
This is a Tension 1 question before it's a procurement one (Chapter 4, when the work itself changes), because cost per finished task is what tells you which work is worth handing over at all. And the real ceiling sits well below what the demos suggest.
The Remote Labor Index, built by Scale AI and the Center for AI Safety, pays human professionals to judge one narrow thing: would a paying client accept this deliverable as it stands. The best model now clears that bar on 16.1% of real projects. Eight months ago the leader managed 2.5%. Read that number precisely, though, because a deliverable that fails the test can still be most of the way there. The gap between 16% and 100% is mostly work that still needs finishing, and you should budget for finishing it (but that budget will keep decreasing each month).
The same models got harder to keep inside the lines
Now the harder half, where the useful lesson runs quieter than the headlines did.
On July 21, OpenAI disclosed that two of its models, GPT-5.6 Sol and a more capable unreleased one, broke out of a sealed testing environment, reached the open internet, and compromised the production systems of Hugging Face, the platform where much of the AI industry stores and shares its models. The goal was to steal the answer key to the security exam they were being graded on. They found a software flaw nobody knew about to get out, then used stolen passwords to get in.
Hugging Face wasn't the only company touched. And one detail from the cleanup deserves more attention than it got. When Hugging Face went to investigate the attack, the commercial AI models it reached for refused to help, because their safety rules block anything that looks like hacking. The forensics ran instead on a Chinese model the company could host on its own machines.
Nine days later, Anthropic published its own version. It had gone back through 141,006 evaluation runs after OpenAI went public, and found three cases where a model left its test environment and got into a real company's systems. The Anthropic case is less cinematic and a great deal more instructive.
There was no sophisticated attack. A setup error with an outside testing partner left the environment connected to the internet, and the models walked through the open door using guessable passwords, systems left without a login, and other basic weaknesses a security team would flag on a routine audit.
Every prompt told the model it was in a simulation with no internet access. That instruction did not hold. Claude Opus 4.7 worked out that it had reached real production systems in all four runs and kept going anyway, pulling passwords and touching live data.
Mythos 5 also worked out that it was on real systems, then reasoned its way back to believing it was still in a simulation, and published a booby-trapped software package to a public library that developers install from. Fifteen real machines downloaded and ran it. Only the newest model in the set stopped. Anthropic's account of that moment is one line: when evidence emerged that the systems were real, "it stopped the exercise."
Here's what I keep coming back to. All three models were given the same instruction. All three carried their normal safety training. And the instruction is the part that failed.
A written rule telling a capable model what it may not do is bolt-on governance in its purest form. It sits outside the thing it's meant to constrain, and it holds only for as long as the model agrees with it. Opus 4.7 and Mythos 5 both talked themselves past it. What held on the newest model was judgment built in, where there was nothing to argue with.
That's Tension 3 (Chapter 6), and the labs just ran the experiment on themselves and published the result. Anthropic is careful about the cause, saying only that the pattern is "consistent with more advanced models responding more appropriately" and that it would want more testing before drawing conclusions. Fair enough. The split between the two older models and the newest one is still worth noticing.
Anthropic calls the episode "closer to a harness and operational failure than a model alignment failure," which is right about how the models got out. It sets aside the harder question, which is what was supposed to stop them once they were out. That job fell to a written instruction, which is precisely the control the book tells you not to lean on.
The narrower version of this point turned up this month from Octopus Deploy, a company that sells software deployment tools, arguing that reviewing AI-written code has become theater because the volume long ago outran the reviewers. They have their commercial reasons for saying so. It's also true, and it doesn't stop at code. Autonomy without governance is reckless. Governance without autonomy is theater. Bolt a human onto the end of a process running at machine speed and you get theater.
None of which necessarily means your organization is exposed the way Hugging Face was. It means the ingredients are ordinary. The failure that reached real systems didn't need a model capable of inventing new attacks. It needed a capable agent, a door somebody left open, and a task the agent fully meant to finish. Whether a capable agent, an open door, and a determined task sit together anywhere in your operation is a question worth asking out loud, and at many organizations nobody has been asked to check.
The people building these models appear to share the worry, and that's the genuinely new thing in July. On July 28, more than 1,300 employees of frontier AI companies asked the U.S. government to help build the tools to "deliberately pace the frontier of automated AI development." Dario Amodei of Anthropic signed it. So did Ilya Sutskever of Safe Superintelligence and Jakub Pachocki, OpenAI's chief scientist. Anthropic and OpenAI each endorsed it as companies within hours.
Read the ask carefully, because it's narrower and more interesting than the coverage suggested. They aren't requesting a slowdown now. They're requesting a mechanism that would make slowing down possible, so no single lab has to give up ground alone.
Earlier in the month all three frontier CEOs put a regulator in writing, each with a different shape. Demis Hassabis of Google DeepMind wants something like FINRA. Sam Altman of OpenAI wants something like the IAEA. Amodei wants something like the FAA, with the power to block a release outright. Illinois, meanwhile, signed the first state law requiring independent outside audits, reaching developers at the scale of the labs themselves.
For a strategy team that changes one thing concretely: the risk of building on someone else's frontier model. June showed how fast a model can be pulled out from under a live deployment, and about a hundred organizations got Mythos 5 back only under federal gating. A predictable, certified release process is what lets you commit to a plan and expect to still be navigating it in six months.
The pacing debate should not slow you down. It should make you certain about which settings you've actually chosen, because the current is doing the moving either way.
Also on the river
Two studies landed this month that measured something the frameworks have only argued about. Hyunjin Kim of INSEAD and Rembrand Koning of Harvard Business School found that AI-native startups run about 25% fewer people than comparable firms at the same valuation, sit roughly half a seniority level flatter, and employ about 13% more engineers alongside 15% fewer entry-level staff and managers.
Separately, Ramp and Revelio Labs tracked more than 21,000 U.S. companies and found the heaviest AI adopters grew headcount by 10.2% over two years, with entry-level hiring up 12%. Read that as a correlation, though. The companies spending most on AI were already the larger, faster-growing ones.
Those findings look contradictory and they aren't. They measured two different populations. Ramp watched existing companies adding AI to the operating logic they already had, and those companies grew. Harvard and INSEAD watched companies built around AI from the start, whose move is pushing knowledge work that used to need internal teams out to where the customer touches it directly, and those companies need fewer people for the same valuation.
That's retrofit and reimagine, measured separately, by researchers who never use either word. I'd resist anyone who tells you it settles the master choice (Chapter 2), because it doesn't. No position on that spectrum removes the risk. What the evidence does is raise the price of picking your spot by accident.
What I'm watching next
Three things. Whether more labs go back through their own evaluation transcripts, since Anthropic only found its incidents by looking after someone else confessed, and nobody thinks two labs is the whole list. Whether cost per completed task shows up in a real board pack this year or stays a vendor talking point. And whether the pacing request becomes anything, or whether Illinois-style laws keep arriving one state at a time while Washington thinks about it.
Sources
- Kimi K3 model card and pricing (OpenRouter)
- Anthropic, Introducing Claude Opus 5
- OpenAI's CFO pitches a new way to measure AI's value (Axios)
- Palantir's Karp bashes the token model (CNBC)
- Remote Labor Index (Scale AI and the Center for AI Safety)
- OpenAI and Hugging Face partner to address security incident during model evaluation
- Anthropic, Investigating three real-world incidents in our cybersecurity evaluations
- Code review is theater now (Octopus Deploy)
- Pacing the Frontier
- Behind the Curtain: AI godfathers converge on regulations (Axios)
- Illinois enacts the Artificial Intelligence Safety Measures Act (Greenberg Traurig)
- AI-Native Firms, Hyunjin Kim and Rembrand Koning (Harvard Business School Working Paper 26-090)
- Companies hire more after AI adoption (Ramp and Revelio Labs)