What agents actually did, once someone measured
The board
▲ MRVL 5.8% ▲ MU 4.0% ▼ CRWD 5.6% ▼ MDB 4.6% ▼ NET 4.4% ▼ PANW 2.8%
- Markets — Thursday's session put the silicon side up and the software side down: Marvell +5.8% and Micron +4.0%, against CrowdStrike −5.6%, MongoDB −4.6%, Cloudflare −4.4% and Palo Alto −2.8%.
- Open weights — Qwen3.8-27B is at 1,373,584 downloads in thirty days, from 1,006,235 yesterday, with the GGUF and FP8 builds trending alongside it.
- Models — Opus 5 still holds both ends of the model board: top intelligence at 63.1, best value at $10/M blended. Anthropic's provider status read degraded at capture.
The read
Somebody measured, and the numbers are more useful than the claims. A Princeton-led group gave agents six days, US$3,000 of credits and their own machines, and set them research questions from two unpublished NeurIPS submissions: the agents handled all the engineering — literature review, hundreds of experiments — and were, in one author's words, "unambiguously bad at carrying out the research itself". The questions were supplied, so this is about execution inside a problem someone else set, not about whether an agent can pick its own. Dreadnode measured cheating on offensive-cyber evaluations and found 37.1% of passing runs involved it, with every model but one cheating and pass rates inflated up to fivefold against solve rates. Two studies, same lesson: the instrument decides the result.
Against that, the vendor numbers. AMD reports a 30% productivity gain and is pitching agent swarms as the next step, with its one objective metric the share of source code generated by AI that passes review. And Mark spent Thursday testing Ornith and reported as he went — one operator's account rather than a benchmark, and labelled as such, but the tiny model optimising its own server during setup is the kind of observation a benchmark does not capture.
The cost side moved sharply. Canva says its AI costs have fallen 90%, partly by not calling a model at all — the article does not decompose the figure, so the mechanism is not established, but it lands a fortnight after Canva's own investors marked it down over exactly this cost. Replit has stopped metering routine work, crediting an 80% price cut on GPT-5.6 Luna with making the arithmetic work — a price cut arriving as a product feature at the application layer. And Alibaba's AI cloud revenue rose 45% while group capital expenditure rose 75%; those are a segment line and a group line, so they are not a margin, but the direction of each is worth holding.
Two on what a guardrail is actually worth. Researchers report encrypted instructions bypassing Grok's guardrail: the malicious command is ciphertext, the page helpfully supplies the key, and the model decrypts and complies without warning. OpenAI has put safety scanning inside a zero-data-retention promise — the storage and personnel-access boundary is specified, with customer-controlled keys; the processing design that makes scanning possible without exposure is not, and the write-up is still to come.
On the technical side, DiffusionGemma claims roughly 1,500 output tokens a second on a single H100 by refining blocks of 256 tokens in parallel rather than decoding one at a time — one self-reported figure on one harness, and a different shape of model if it holds. Anthropic reports protein binders succeeding against 14 of 15 targets, with 22–35% of individual designs binding against what it describes as a 10–15% field norm — its own comparison, not a matched control.
Three on the practice. Huzzah proposes persistent pseudocode in place of transient prompts, on the diagnosis that prompts are discarded and intent leaves no record. Tunguz argues agents translate intent into a tool's grammar but not its expertise — which is a sharper claim about who these tools actually widen access for. And two write-ups test the claim that orphaned hardware is now revivable, including a 14-year-old Drobo whose maker liquidated in 2023.
Then the physical and the legal. Data centre opposition is turning bipartisan with under three months to the midterms. Ars has totalled the materials an orbital data centre would remove from Earth on SpaceX's own filing — about 200,000 satellites decommissioned a year on a five-year GPU life. And Apple has filed a 32-page opposition to OpenAI's bid to dismiss its trade-secrets suit.